{"schemaVersion":"jobsearcher.job.v1","id":"e571c9933a79bce620352edc","url":"https://jobsearcher.com/jobs/e571c9933a79bce620352edc","canonicalUrl":"https://jobsearcher.com/jobs/e571c9933a79bce620352edc","title":"ML Infra Engineer - Supercomputing","description":"Physical Intelligence builds general-purpose AI for the physical world. Training our models requires orchestrating thousands of accelerators across a heterogeneous fleet of GPU and TPU clusters — spanning different hardware generations, cloud providers, and cluster topologies.\nToday, researchers often need to know which cluster to target, what resources are available, and how to configure their jobs accordingly. That doesn't scale. We need a scheduling and compute layer that makes the right placement decision automatically — routing jobs to the best cluster based on availability, hardware fit, cost, and priority — so researchers can focus entirely on the science.\nThis role owns that problem end-to-end: the scheduling systems, the placement logic, the cluster management layer, and the operational tooling that keeps it all running.\nThis is not cloud DevOps. It's not about standing up clusters and walking away. It's a systems role for people who care about intelligent resource allocation, utilization, fault tolerance, and making large-scale distributed training seamless.\nThe Team\nThe ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. You will work closely with ML Infra (training systems), data platform, and research teams to ensure compute scheduling is never the bottleneck.\nIn This Role You Will\nOwn Intelligent Job Scheduling and Placement: Design and build multi-tenant scheduling systems that automatically place training jobs on the best available cluster based on hardware requirements, topology, availability, cost, and priority. Support fair resource sharing across teams and projects with quota management, priority tiers, and preemption policies. Abstract away cluster differences so researchers submit jobs without needing to know where they will land.\nScale Multi-cluster Orchestration: Build the control plane that manages the job lifecycle across diverse clusters (mixed GPU/TPU, multi-generation hardware, on-prem/cloud) and enables seamless job migration, failover, and re-scheduling.\nOptimize Accelerator Utilization and Efficiency: Monitor and optimize GPU/TPU utilization across the entire fleet. Implement priority, preemption, queueing, and fairness policies that balance research velocity with cost efficiency.\nEnsure Scaling and Stability: Implement fault detection, automatic recovery, and resilience for long-running multi-node training jobs. Manage health checking, node management, and scaling to thousands of accelerators.\nSupport Inference and Robot Deployment: Extend scheduling and orchestration to inference workloads, including deploying models to edge devices on physical robots.\nEnhance Observability and Developer Experience: Build the dashboards, alerting, SLOs, and debugging tools necessary for researchers to understand job status and for the team to ensure high scheduling quality and cluster reliability.\nWhat We Hope You’ll Bring\nWe’re intentionally flexible on exact background, but strong candidates usually have:\nStrong software engineering fundamentals\nExperience building or operating job scheduling / resource management systems at scale\nExperience with large-scale compute clusters (GPU and/or TPU)\nFamiliarity with schedulers and orchestration systems (SLURM, Kubernetes, GKE, K3S, or internal equivalents)\nComfort reasoning about resource allocation, bin-packing, priority scheduling, and multi-tenancy\nUnderstanding of how ML training workloads behave — long-running, multi-node, sensitive to stragglers, topology-dependent\nA bias toward owning systems end-to-end, from design to operation\nEnjoy working closely with researchers and unblocking fast-moving projects\nBonus Points If You Have\nExperience building multi-cluster or federated scheduling systems\nExperience with TPU infrastructure (GCP TPU slices, Multislice, GKE)\nBackground in cluster resource managers (Borg, YARN, Mesos, or custom schedulers)\nLinux systems engineering, networking, and infrastructure-as-code\nNCCL/collective communication and topology-aware placement\nExperience with capacity planning and cloud cost optimization at scale\nFamiliarity with JAX, PyTorch, or similar ML frameworks at the runtime/systems level\nIn this role you will help scale and optimize our training systems and core model code. You’ll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You’ll work closely with researchers and model engineers to translate ideas into experiments—and those experiments into production training runs.\nThis is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.\nThe Team\nThe ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.\nIn This Role You Will\nOwn training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.\nScale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction.\nOptimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.\nEnable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.\nManage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.\nPartner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale.\nContribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.\nWhat We Hope You’ll Bring\nStrong software engineering fundamentals and experience building ML training infrastructure or internal platforms.\nHands-on large-scale training experience in JAX (preferred), PyTorch.\nFamiliarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.\nExperience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS).\nAbility to debug and optimize performance bottlenecks across the training stack.\nStrong cross-functional communication and ownership mindset.\nBonus Points If You Have\nDeep ML systems background (e.g., training compilers, runtime optimization, custom kernels).\nExperience operating close to hardware (GPU/TPU performance tuning).\nBackground in robotics, multimodal models, or large-scale foundation models.\nExperience designing abstractions that balance researcher flexibility with system reliability.\nPursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.","company":"Physical Intelligence","rawCompany":"physical intelligence","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-04-12T19:37:32.852Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"ML Infra Engineer - Supercomputing","description":"Physical Intelligence builds general-purpose AI for the physical world. Training our models requires orchestrating thousands of accelerators across a heterogeneous fleet of GPU and TPU clusters — spanning different hardware generations, cloud providers, and cluster topologies.\nToday, researchers often need to know which cluster to target, what resources are available, and how to configure their jobs accordingly. That doesn't scale. We need a scheduling and compute layer that makes the right placement decision automatically — routing jobs to the best cluster based on availability, hardware fit, cost, and priority — so researchers can focus entirely on the science.\nThis role owns that problem end-to-end: the scheduling systems, the placement logic, the cluster management layer, and the operational tooling that keeps it all running.\nThis is not cloud DevOps. It's not about standing up clusters and walking away. It's a systems role for people who care about intelligent resource allocation, utilization, fault tolerance, and making large-scale distributed training seamless.\nThe Team\nThe ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. You will work closely with ML Infra (training systems), data platform, and research teams to ensure compute scheduling is never the bottleneck.\nIn This Role You Will\nOwn Intelligent Job Scheduling and Placement: Design and build multi-tenant scheduling systems that automatically place training jobs on the best available cluster based on hardware requirements, topology, availability, cost, and priority. Support fair resource sharing across teams and projects with quota management, priority tiers, and preemption policies. Abstract away cluster differences so researchers submit jobs without needing to know where they will land.\nScale Multi-cluster Orchestration: Build the control plane that manages the job lifecycle across diverse clusters (mixed GPU/TPU, multi-generation hardware, on-prem/cloud) and enables seamless job migration, failover, and re-scheduling.\nOptimize Accelerator Utilization and Efficiency: Monitor and optimize GPU/TPU utilization across the entire fleet. Implement priority, preemption, queueing, and fairness policies that balance research velocity with cost efficiency.\nEnsure Scaling and Stability: Implement fault detection, automatic recovery, and resilience for long-running multi-node training jobs. Manage health checking, node management, and scaling to thousands of accelerators.\nSupport Inference and Robot Deployment: Extend scheduling and orchestration to inference workloads, including deploying models to edge devices on physical robots.\nEnhance Observability and Developer Experience: Build the dashboards, alerting, SLOs, and debugging tools necessary for researchers to understand job status and for the team to ensure high scheduling quality and cluster reliability.\nWhat We Hope You’ll Bring\nWe’re intentionally flexible on exact background, but strong candidates usually have:\nStrong software engineering fundamentals\nExperience building or operating job scheduling / resource management systems at scale\nExperience with large-scale compute clusters (GPU and/or TPU)\nFamiliarity with schedulers and orchestration systems (SLURM, Kubernetes, GKE, K3S, or internal equivalents)\nComfort reasoning about resource allocation, bin-packing, priority scheduling, and multi-tenancy\nUnderstanding of how ML training workloads behave — long-running, multi-node, sensitive to stragglers, topology-dependent\nA bias toward owning systems end-to-end, from design to operation\nEnjoy working closely with researchers and unblocking fast-moving projects\nBonus Points If You Have\nExperience building multi-cluster or federated scheduling systems\nExperience with TPU infrastructure (GCP TPU slices, Multislice, GKE)\nBackground in cluster resource managers (Borg, YARN, Mesos, or custom schedulers)\nLinux systems engineering, networking, and infrastructure-as-code\nNCCL/collective communication and topology-aware placement\nExperience with capacity planning and cloud cost optimization at scale\nFamiliarity with JAX, PyTorch, or similar ML frameworks at the runtime/systems level\nIn this role you will help scale and optimize our training systems and core model code. You’ll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You’ll work closely with researchers and model engineers to translate ideas into experiments—and those experiments into production training runs.\nThis is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.\nThe Team\nThe ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.\nIn This Role You Will\nOwn training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.\nScale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction.\nOptimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.\nEnable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.\nManage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.\nPartner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale.\nContribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.\nWhat We Hope You’ll Bring\nStrong software engineering fundamentals and experience building ML training infrastructure or internal platforms.\nHands-on large-scale training experience in JAX (preferred), PyTorch.\nFamiliarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.\nExperience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS).\nAbility to debug and optimize performance bottlenecks across the training stack.\nStrong cross-functional communication and ownership mindset.\nBonus Points If You Have\nDeep ML systems background (e.g., training compilers, runtime optimization, custom kernels).\nExperience operating close to hardware (GPU/TPU performance tuning).\nBackground in robotics, multimodal models, or large-scale foundation models.\nExperience designing abstractions that balance researcher flexibility with system reliability.\nPursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.","datePosted":"2026-04-12T19:37:32.852Z","dateModified":"2026-04-12T19:37:32.852Z","hiringOrganization":{"@type":"Organization","name":"Physical Intelligence","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e571c9933a79bce620352edc"},"url":"https://jobsearcher.com/jobs/e571c9933a79bce620352edc"}}