Software Systems Engineering
Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence. We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals. To be considered for an interview, please make sure your application is full in line with the job specs as found below.We are scientists, engineers, and builders who've created some of the most widely used AI products, including ChatGPT and , open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything. We're looking for generalist infrastructure and systems engineers to help build the systems that power our foundation models and the internal teams on research and product development to be able to create the models and ship the products powered by our models. You'll work across the full technical stack, solving complex distributed systems problems and building robust, scalable platforms. You'll work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights across our models, products, and data assets. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. We continuously review applications and reach out to applicants as new opportunities open. You may also find that we put up postings for singular roles for separate, project or team specific needs. We interview generally, but during project selection we'll take into account your interests and experience alongside organizational needs. This flexible approach allows us to match talented engineers with the infrastructure teams where they'll have the greatest impact and growth potential. Core Infrastructure: We support teams that train, research, and ultimately serve AI models and build the underlying infrastructure for the clusters to reliably and safely train frontier models. Examples might include building systems and running large Kubernetes clusters with GPU workloads, or building infrastructure to support Tinker. Data Infrastructure: We build and maintain the data systems for our research and products. You'll design and optimize data pipelines using tools like Spark and other modern data infrastructure technologies. You'll build scalable, reliable, data infrastructure while embedding governance best practices. We care deeply about research and engineering productivity and our ability to continue shipping quickly. Bachelor's degree or equivalent experience in computer science, engineering, or similar. Proficiency in at least one backend language (we use Python or Rust). Comfort operating across the stack and owning projects end-to-end. A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships. Strong debugging across application, OS, and network layers. Proficiency in Python or Rust (or similar), containers, and modern CI. Experience with Kubernetes, controllers/operators, or performance profiling. Familiarity with GPU/ML workflows or large-scale data/eval pipelines. Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together. xhyhwjd Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed. As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.