JOBSEARCHER

RL Systems Engineer: Large-Scale GPU Training

Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal intervention. You will design and optimize training pipelines across GPUs, build observability tooling, and collaborate on post-training capabilities for production deployments. #J-18808-Ljbffr