Lead Large-Scale GPU Cluster Engineer for AI Research
Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.
The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.
#J-18808-Ljbffr