JOBSEARCHER

SRE - Compute & Hyperscale GPU Fleet Reliability

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling. The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation. #J-18808-Ljbffr