JOBSEARCHER

AI Infrastructure Engineer

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

AI Infrastructure Engineer – Distributed ML SystemsHybrid (San Francisco, CA)AI Infrastructure | Distributed Systems | MLOpsAbout the RoleWe’re looking for an AI Infrastructure Engineer to build and scale the systems that power large-scale machine learning workloads.You’ll work on distributed training, model serving, and high-performance ML infrastructure, enabling AI teams to train and deploy models efficiently at scale.Key ResponsibilitiesDesign and maintain distributed ML training infrastructureBuild scalable model serving systemsOptimise GPU utilisation and training performanceDevelop tools for experiment tracking and model lifecycle managementWork closely with AI/ML teams to improve workflowsRequired Skills & Experience5+ years in backend, infrastructure, or ML systems engineeringStrong Python + systems programming experienceExperience with distributed systems and cloud platformsExperience with Kubernetes and containerised environmentsFamiliarity with ML frameworks and workflowsNice to HaveExperience with distributed training (PyTorch Distributed, Ray, etc.)GPU optimisation and high-performance computingExperience with ML platforms or internal tooling