JOBSEARCHER

GPU Reliability Engineer Systems & Firmware

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

Thinking Machines Lab is hiring an engineer to ensure the reliability of our GPU supercomputing fleet. You will own the seam between hardware, firmware, and OS, diagnosing issues to root cause and coordinating fixes with vendors so researchers can scale confidently. This evergreen role seeks a proactive engineer who thrives in a highly collaborative environment and can drive cross-stack improvements across kernels, drivers, and hardware health signals. #J-18808-Ljbffr