Training - Runtime Foundations Engineer
Runtime - San Francisco
About the Team
Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We build robustance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, accelerating progress towards AGI.
Within Training Runtimeance, stability, and observability.
Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs.
About the Role
As a Training Runtime: Runtime Foundations Engineer, you will work on the software that ties thousands of computers together and exposes them as a unified system.
This system has to serve individual researchers running multiple parallel experimentsance throughout.
You will work primarily in Codex + Rustance, correctness, resilience, and scalability.
Working at this scale and at the frontier of AI development poses novel challenges. We support many different types of research, and the problems you will be working on are highly ambiguous and require a superb ability to jump into novel domains, understand user problems, and demonstrate both strong design judgment and proficient execution to advance our programs.
We’re looking for people who love optimizing an end-to-end platformance across our supercomputers. We’re looking for engineers excited by the rapid pace of responding to the dynamic and evolving needs of our training runtime and compute stack.
This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
In this role, you will:
Work across our Python and Rust stack
Design, build, and maintain software to orchestrate and monitor machine learning workloads on our largest supercomputers
Profile and optimize our software stack to support computation orchestration at frontier scale
Improve reliability, observability, and fault tolerance for long-running jobs
Debug complex distributed systems issues across large clusters
Respond to the changing shapes and needs of the ML systems to enable our researchers
You might thrive in this role if you:
Have experience developing distributed systems
Enjoy understanding how large systems behave and fail at scale
Love being both a developer and an operator
Care deeply about performance, correctness, and reliability
Have strong software engineering skills and are proficient in Rust or another systems programming language (e.g. C++)
Have solid Linux knowledgeance analysis, and memory profiling
Are comfortable and experienced in developing asynchronous and concurrent systems
Like high-ownership environments with light process and strong engineering agency
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core the full spectrum of humanity.
We are an equal opportunity employeration, or other applicable legally protected characteristic.
Background checks for applicants will be administered in accordance with applicable lawation technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities.
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Compensation
$266K – $385K + Offers Equity