JOBSEARCHER

Infrastructure Engineer

CalanceCosta Mesa, CAL5 SeniorSeptember 19th, 2026
Senior AI Infrastructure Engineer – Physical InfrastructureLocation: Costa Mesa, CA | OnsiteType: Full-Time / PermanentWe’re seeking a Senior AI Infrastructure Engineer to own large-scale GPU infrastructure supporting AI training and inference.Key Responsibilities:Rack, stack, cable, and bring up H200/B200/B300/NVL72 GPU systems.Build and optimize NVLink, InfiniBand, RoCE, and Spectrum-X networks.Deploy and manage VAST, DDN, Weka, Lustre or similar high-performance storage.Manage Kubernetes, Run:AI/Ray for GPU scheduling and multi-tenancy.Automate cluster deployment, firmware/driver management, and infrastructure configuration.Monitor and troubleshoot GPU Xid, NCCL, networking, and hardware failures.Build automated fault detection, observability, and self-healing capabilities.Must Have:10+ years infrastructure/HPC/datacenter engineering experience.Hands-on large-scale NVIDIA GPU cluster experience.Strong InfiniBand/RoCE/NVLink experience.Kubernetes + GPU orchestration experience.High-performance parallel storage experience.Strong automation/IaC skills.Able to perform physical datacenter work and lift 50+ lbs.Eligible for U.S. Top Secret clearance.