Staff Engineer, Large-Scale GPU Pre-Training Systems
Magic AI, Inc. is seeking a Member of Technical Staff for Pre-training Systems to design and operate distributed infrastructure that trains long-context models at scale.
The role focuses on large-scale model training across thousands of GPUs, ensuring performance, reliability, and reproducibility in extreme-scale environments. You will own systems handling memory pressure, cross-device communication, fault-tolerant jobs, and efficient sequence packing, collaborating with Kernels and Research to
#J-18808-Ljbffr