Machine Learning Infrastructure Engineer (Modeling)
The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs
Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization
Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments
Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale
Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics#J-18808-Ljbffr