GPU Reliability Engineer Systems & Firmware
Thinking Machines Lab is hiring an engineer to ensure the reliability of our GPU supercomputing fleet. You will own the seam between hardware, firmware, and OS, diagnosing issues to root cause and coordinating fixes with vendors so researchers can scale confidently.
This evergreen role seeks a proactive engineer who thrives in a highly collaborative environment and can drive cross-stack improvements across kernels, drivers, and hardware health signals.
#J-18808-Ljbffr