Staff Engineer: Distributed ML Training Systems
reflectionai seeks a senior engineer to build and scale distributed training systems powering frontier model pre-training in San Francisco. You will partner with research teams to run large-scale foundation model training across thousands of GPUs, designing infrastructure for efficient, production-ready workflows.
You will optimize throughput, memory, and GPU utilization while debugging complex distributed pipelines and collaborating with researchers to push novel training techniques into
#J-18808-Ljbffr