RL Systems Engineer: Large-Scale GPU Training
Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal intervention.
You will design and optimize training pipelines across GPUs, build observability tooling, and collaborate on post-training capabilities for production deployments.
#J-18808-Ljbffr