JOBSEARCHER

GPU Reliability Engineer

EvonaMillbrae, CAL6 LeadOctober 3rd, 2026
EVONA is partnering with a fast-growing space technology company developing high-performance compute infrastructure for Low Earth Orbit. They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms. When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously. What You’ll Do Develop fault injection, detection and recovery testing Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink Characterise fault propagation across hardware, firmware and OS layers Validate ECC, MCA/MCI and system recovery mechanisms Work with BMC, IPMI, Redfish and MCTP/PLDM Partner directly with silicon vendors on failure analysis and root cause Build Python automation for testing and log analysis What We’re Looking For 5+ years in hardware validation, silicon validation or platform reliability Strong CPU/GPU architecture knowledge Experience with DDR/HBM and server-class compute systems Deep understanding of RAS, ECC, fault containment and recovery Experience with PCIe, NVLink and/or XGMI Strong Python or equivalent scripting skills Why Join? You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space. US export control requirements apply to this position. #J-18808-Ljbffr