{"schemaVersion":"jobsearcher.job.v1","id":"f8877193a67ec76f7e4e199d","url":"https://jobsearcher.com/jobs/f8877193a67ec76f7e4e199d","canonicalUrl":"https://jobsearcher.com/jobs/f8877193a67ec76f7e4e199d","title":"Senior Software Engineer, DGX Cloud Production Engineering","description":"Overview\nIn this role you will help build and operate automation and tooling for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments. You will develop services for provisioning, validation, upgrades, monitoring, and lifecycle operations, while advancing Day 0–2 workflows and GitOps-driven production readiness. You’ll reduce manual touches through automation and contribute to on-call and incident response. You’ll collaborate with platform, storage, networking, and security teams to deliver reliable, scalable infrastructure. This is a hands-on role at the heart of DGX Cloud’s GPU infrastructure and production operations.\n\nCompensation / Benefitsequitybenefitsremote work options\nResponsibilitiesBuild and operate automation for large-scale GPU clusters across cloud partners and on-premDevelop tools for provisioning, validation, upgrades, monitoring, repair, and lifecycle operationsImprove Day 0/1/2 workflows for cluster bringup and production handoffReduce manual production touches via APIs, GitOps, automation, and agent-assisted workflowsParticipate in on-call, incident response, and durable follow-up workPartner with platform, storage, networking, and security teams to make infrastructure production-ready\nKey requirements8+ years of experience in production infrastructureStrong programming skills in Python, Go, or similarExperience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automationAbility to troubleshoot distributed systems in productionClear communication and cross-team collaborationBS/MS in Computer Science or equivalent experienceStrong communicationCross-team collaborationProblem-solving under pressurePythonGoLinux","company":"NVIDIA","rawCompany":"nvidia","city":"Bloomington","state":"IL","isRemote":false,"isActive":false,"createdAt":"2026-09-22T03:40:37.598Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Software Engineer, DGX Cloud Production Engineering","description":"Overview\nIn this role you will help build and operate automation and tooling for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments. You will develop services for provisioning, validation, upgrades, monitoring, and lifecycle operations, while advancing Day 0–2 workflows and GitOps-driven production readiness. You’ll reduce manual touches through automation and contribute to on-call and incident response. You’ll collaborate with platform, storage, networking, and security teams to deliver reliable, scalable infrastructure. This is a hands-on role at the heart of DGX Cloud’s GPU infrastructure and production operations.\n\nCompensation / Benefitsequitybenefitsremote work options\nResponsibilitiesBuild and operate automation for large-scale GPU clusters across cloud partners and on-premDevelop tools for provisioning, validation, upgrades, monitoring, repair, and lifecycle operationsImprove Day 0/1/2 workflows for cluster bringup and production handoffReduce manual production touches via APIs, GitOps, automation, and agent-assisted workflowsParticipate in on-call, incident response, and durable follow-up workPartner with platform, storage, networking, and security teams to make infrastructure production-ready\nKey requirements8+ years of experience in production infrastructureStrong programming skills in Python, Go, or similarExperience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automationAbility to troubleshoot distributed systems in productionClear communication and cross-team collaborationBS/MS in Computer Science or equivalent experienceStrong communicationCross-team collaborationProblem-solving under pressurePythonGoLinux","datePosted":"2026-09-22T03:40:37.598Z","dateModified":"2026-09-22T03:40:37.598Z","hiringOrganization":{"@type":"Organization","name":"NVIDIA","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bloomington","addressRegion":"IL","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f8877193a67ec76f7e4e199d"},"url":"https://jobsearcher.com/jobs/f8877193a67ec76f7e4e199d"}}