Machine Learning Infrastructure
A well-funded Series A AI infrastructure company is looking for an exceptional distributed-systems/ML Infra engineer to help build the orchestration layer for the future of AI compute.You’ll design and operate systems that schedule, route, and coordinate production AI workloads across thousands of nodes and heterogeneous hardware.We’d like to hear from engineers who have personally built and owned:Production schedulers or orchestration platformsDistributed control planesResource-management or queueing systemsKubernetes-adjacent systems - not simply operated clustersLarge-scale ML serving or distributed-compute infrastructureYou should be comfortable reasoning about concurrency, coordination, fault tolerance, failure recovery, and production reliability. Experience with Go, C++, Python, RPC, or asynchronous systems is especially relevant.This isn’t a standard cloud-platform role. We’re looking for a builder who can clearly explain what they personally designed, shipped, scaled, and operated.You’ll have the opportunity to shape both the foundational infrastructure and engineering culture of a company building for the next generation of AI.