{"schemaVersion":"jobsearcher.job.v1","id":"6b11ed01cb315fd5b9d89ca2","url":"https://jobsearcher.com/jobs/6b11ed01cb315fd5b9d89ca2","canonicalUrl":"https://jobsearcher.com/jobs/6b11ed01cb315fd5b9d89ca2","title":"Principal Deployment Engineer","description":"Principal Deployment Engineer – GPU Supercluster Bringup\nAbout Us\nWe are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.\nAs a startup, we move fast, operate with ownership, and expect technical leaders to define standards—not just follow them.\nThe Role\nWe are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.\nYou will own the full lifecycle of deployment—from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.\nThis role combines deep hands-on expertise with system-level thinking and cross-functional leadership.\nWhat You'll Do\nEnd-to-End Supercluster Bringup Ownership\nDefine the technical standards for node, rack, and full-cluster bringup.\n\nLead large-scale GPU cluster deployments (multi-rack, multi-pod environments). Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads. Establish cluster-level acceptance criteria and validation frameworks.\nPerformance & Fabric Architecture\nTune and validate NCCL, RDMA, GPUDirect, and collective operations at scale.\n\nIdentify and eliminate performance bottlenecks across hardware, topology, and firmware layers. Drive congestion control and fabric optimization strategies. Define performance benchmarking methodology for AI training workloads.\nDeployment Strategy & Scalability\nDesign repeatable deployment models for multi-site expansion.\n\nBuild automation frameworks for provisioning and cluster validation. Establish deployment SLAs, quality gates, and operational readiness standards. Reduce time-to-capacity while increasing reliability.\nTechnical Leadership\nServe as the escalation point for complex bringup and performance issues.\n\nMentor senior engineers and shape infrastructure best practices. Influence hardware selection, rack topology, and data center design decisions. Partner with executive leadership on infrastructure scaling strategy.\nWhat We're Looking For\nRequired\n10+ years of experience in large-scale infrastructure or HPC environments.\n\nProven experience bringing up large GPU clusters (hundreds+ GPUs). Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics). Strong understanding of server architecture (PCIe, NUMA, memory hierarchy). Experience debugging performance issues across compute and network layers. Strong automation and systems-level thinking.\nStrongly Preferred\nExperience scaling AI training clusters for frontier models.\n\nExperience with liquid cooling or ultra-high-density deployments. Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF). Experience defining infrastructure standards in a fast-growing organization.\nWhat Success Looks Like\nSuperclusters are brought online quickly, predictably, and at peak performance.\n\nDeployment processes scale from first cluster to multi-site expansion. Infrastructure becomes a competitive advantage.\nYou define the technical blueprint for how we scale AI infrastructure.\nFor information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.","company":"Nscale","rawCompany":"nscale","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-03T15:25:45.692Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1241.00","title":"Computer Network Architects","slug":"computer-network-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal Deployment Engineer","description":"Principal Deployment Engineer – GPU Supercluster Bringup\nAbout Us\nWe are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.\nAs a startup, we move fast, operate with ownership, and expect technical leaders to define standards—not just follow them.\nThe Role\nWe are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.\nYou will own the full lifecycle of deployment—from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.\nThis role combines deep hands-on expertise with system-level thinking and cross-functional leadership.\nWhat You'll Do\nEnd-to-End Supercluster Bringup Ownership\nDefine the technical standards for node, rack, and full-cluster bringup.\n\nLead large-scale GPU cluster deployments (multi-rack, multi-pod environments). Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads. Establish cluster-level acceptance criteria and validation frameworks.\nPerformance & Fabric Architecture\nTune and validate NCCL, RDMA, GPUDirect, and collective operations at scale.\n\nIdentify and eliminate performance bottlenecks across hardware, topology, and firmware layers. Drive congestion control and fabric optimization strategies. Define performance benchmarking methodology for AI training workloads.\nDeployment Strategy & Scalability\nDesign repeatable deployment models for multi-site expansion.\n\nBuild automation frameworks for provisioning and cluster validation. Establish deployment SLAs, quality gates, and operational readiness standards. Reduce time-to-capacity while increasing reliability.\nTechnical Leadership\nServe as the escalation point for complex bringup and performance issues.\n\nMentor senior engineers and shape infrastructure best practices. Influence hardware selection, rack topology, and data center design decisions. Partner with executive leadership on infrastructure scaling strategy.\nWhat We're Looking For\nRequired\n10+ years of experience in large-scale infrastructure or HPC environments.\n\nProven experience bringing up large GPU clusters (hundreds+ GPUs). Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics). Strong understanding of server architecture (PCIe, NUMA, memory hierarchy). Experience debugging performance issues across compute and network layers. Strong automation and systems-level thinking.\nStrongly Preferred\nExperience scaling AI training clusters for frontier models.\n\nExperience with liquid cooling or ultra-high-density deployments. Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF). Experience defining infrastructure standards in a fast-growing organization.\nWhat Success Looks Like\nSuperclusters are brought online quickly, predictably, and at peak performance.\n\nDeployment processes scale from first cluster to multi-site expansion. Infrastructure becomes a competitive advantage.\nYou define the technical blueprint for how we scale AI infrastructure.\nFor information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.","datePosted":"2026-08-03T15:25:45.692Z","dateModified":"2026-08-03T15:25:45.692Z","hiringOrganization":{"@type":"Organization","name":"Nscale","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6b11ed01cb315fd5b9d89ca2"},"url":"https://jobsearcher.com/jobs/6b11ed01cb315fd5b9d89ca2"}}