{"schemaVersion":"jobsearcher.job.v1","id":"867d67418b188f349a5823fd","url":"https://jobsearcher.com/jobs/867d67418b188f349a5823fd","canonicalUrl":"https://jobsearcher.com/jobs/867d67418b188f349a5823fd","title":"Staff Software Engineer - Compute","description":"Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.\n\nIf you'd like to build the world's best AI cloud, join us.\n\n*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.\n\nAbout the Role\n\nAs a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.\n\nThe position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.\n\nWhat You'll Do\n\nWe are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:\n\nDesigning and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.\n\nGuide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.\n\nGuide design of compute platform multi-tenant security model\n\nProvide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.\n\nCollaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.\n\nWork with customers on translating vague customer technical requirements into concrete engineering deliverables.\n\nSet engineering standards and lead design reviews for mission-critical cloud software at scale.\n\nWho You are\n\n10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.\n\nDeep expertise in durable execution models and distributed systems used in cloud-service provisioning.\n\nBasic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.\n\nProven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.\n\nProven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).\n\nProficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.\n\nNice to Have\n\nKnowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).\n\nKnowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)\n\nKnowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).\n\nExperience with Cloud Service Provider Kubernetes offerings.\n\nKnowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).\n\nSalary Range Information\n\nThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.\n\nAbout Lambda\n\nFounded in 2012, with 500+ employees, and growing fast\n\nOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove\n\nWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG\n\nOur values are publicly available: https://lambda.ai/careers\n\nWe offer generous cash & equity compensation\n\nHealth, dental, and vision coverage for you and your dependents\n\nWellness and commuter stipends for select roles\n\n401k Plan with 2% company match (USA employees)\n\nFlexible paid time off plan that we all actually use\n\nEqual Opportunity Employer\n\nLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.\n\nCompensation Range: $314K - $465K","company":"Lambda","rawCompany":"lambda","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-19T11:52:34.947Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.00","title":"Computer Occupations, All Other","slug":"computer-occupations-all-other"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Staff Software Engineer - Compute","description":"Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.\n\nIf you'd like to build the world's best AI cloud, join us.\n\n*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.\n\nAbout the Role\n\nAs a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.\n\nThe position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.\n\nWhat You'll Do\n\nWe are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:\n\nDesigning and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.\n\nGuide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.\n\nGuide design of compute platform multi-tenant security model\n\nProvide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.\n\nCollaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.\n\nWork with customers on translating vague customer technical requirements into concrete engineering deliverables.\n\nSet engineering standards and lead design reviews for mission-critical cloud software at scale.\n\nWho You are\n\n10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.\n\nDeep expertise in durable execution models and distributed systems used in cloud-service provisioning.\n\nBasic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.\n\nProven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.\n\nProven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).\n\nProficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.\n\nNice to Have\n\nKnowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).\n\nKnowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)\n\nKnowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).\n\nExperience with Cloud Service Provider Kubernetes offerings.\n\nKnowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).\n\nSalary Range Information\n\nThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.\n\nAbout Lambda\n\nFounded in 2012, with 500+ employees, and growing fast\n\nOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove\n\nWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG\n\nOur values are publicly available: https://lambda.ai/careers\n\nWe offer generous cash & equity compensation\n\nHealth, dental, and vision coverage for you and your dependents\n\nWellness and commuter stipends for select roles\n\n401k Plan with 2% company match (USA employees)\n\nFlexible paid time off plan that we all actually use\n\nEqual Opportunity Employer\n\nLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.\n\nCompensation Range: $314K - $465K","datePosted":"2026-08-19T11:52:34.947Z","dateModified":"2026-08-19T11:52:34.947Z","hiringOrganization":{"@type":"Organization","name":"Lambda","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"867d67418b188f349a5823fd"},"url":"https://jobsearcher.com/jobs/867d67418b188f349a5823fd"}}