{"schemaVersion":"jobsearcher.job.v1","id":"da12602156cf2e2c43746310","url":"https://jobsearcher.com/jobs/da12602156cf2e2c43746310","canonicalUrl":"https://jobsearcher.com/jobs/da12602156cf2e2c43746310","title":"Principal Engineer, Inference Cloud","description":"Overview\nAs Principal Engineer for the Inference Cloud Platform, you translate ambitious scale and reliability goals into a long-term architecture. You will own multi-region topology, failure domains, and system evolution while delivering fast, available inference services. You’ll lead across ML, Product, and Infrastructure teams to solve high-leverage platform problems and drive execution when specs are unclear. Expect to shape platform decisions, production code on critical paths, and strong operational rigor in a rapidly growing setting.\n\nResponsibilitiesDefine and prioritize core platform problems and trade-offs for the Inference Cloud PlatformSet long-term technical direction, topology, and system boundaries for multi-region servicesArchitect reliable, low-latency systems with active-active failover and graceful degradationContribute production code on critical paths and review designs with long-term operational impactLead production issues, drive observability, incident response, and capacity planningInfluence cross-functional teams on API design, deployment strategy, and shared infrastructure decisionsMentor teams to improve technical decision-making and engineering standards\nKey requirements10+ years of software engineering experienceStrong distributed systems and cloud infrastructure backgroundProven ability to design and operate highly available, latency-sensitive systems at scaleExperience optimizing latency, throughput, and efficiency in high-QPS workloadsProduction-ready coding ability in Go, C++, or PythonExperience with observability practices (metrics/logging/tracing) and SLO/SLA-driven operationsAbility to influence senior engineers and cross-functional partners through credibility and judgmentExperience with ML inference infrastructure, model serving, or GPU-accelerated workloads is a plusstrong technical judgmentexcellent communicationleadership and mentorshipdistributed systems architecture in cloud environmentsnetworking, compute orchestration, container platformsmulti-region production services","company":"Cerebras Systems","rawCompany":"cerebras systems","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T04:12:35.985Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal Engineer, Inference Cloud","description":"Overview\nAs Principal Engineer for the Inference Cloud Platform, you translate ambitious scale and reliability goals into a long-term architecture. You will own multi-region topology, failure domains, and system evolution while delivering fast, available inference services. You’ll lead across ML, Product, and Infrastructure teams to solve high-leverage platform problems and drive execution when specs are unclear. Expect to shape platform decisions, production code on critical paths, and strong operational rigor in a rapidly growing setting.\n\nResponsibilitiesDefine and prioritize core platform problems and trade-offs for the Inference Cloud PlatformSet long-term technical direction, topology, and system boundaries for multi-region servicesArchitect reliable, low-latency systems with active-active failover and graceful degradationContribute production code on critical paths and review designs with long-term operational impactLead production issues, drive observability, incident response, and capacity planningInfluence cross-functional teams on API design, deployment strategy, and shared infrastructure decisionsMentor teams to improve technical decision-making and engineering standards\nKey requirements10+ years of software engineering experienceStrong distributed systems and cloud infrastructure backgroundProven ability to design and operate highly available, latency-sensitive systems at scaleExperience optimizing latency, throughput, and efficiency in high-QPS workloadsProduction-ready coding ability in Go, C++, or PythonExperience with observability practices (metrics/logging/tracing) and SLO/SLA-driven operationsAbility to influence senior engineers and cross-functional partners through credibility and judgmentExperience with ML inference infrastructure, model serving, or GPU-accelerated workloads is a plusstrong technical judgmentexcellent communicationleadership and mentorshipdistributed systems architecture in cloud environmentsnetworking, compute orchestration, container platformsmulti-region production services","datePosted":"2026-09-15T04:12:35.985Z","dateModified":"2026-09-15T04:12:35.985Z","hiringOrganization":{"@type":"Organization","name":"Cerebras Systems","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"da12602156cf2e2c43746310"},"url":"https://jobsearcher.com/jobs/da12602156cf2e2c43746310"}}