GPU Platform Reliability Engineer
Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production.
You will automate deployment debugging, create a scalable firmware telemetry stack, and define the qualification process for new GPUs onboarded to our platform. Join Beam at the ground floor of a rapidly growing startup.
#J-18808-Ljbffr