Engineering Manager, GPU Infrastructure
Industry: Hyperscale and AI Data Center and Cloud ComputingLocation: Remote (US, Pacific Time Zone)Employment Type: Full-TimeReporting to: VP, OperationsPosition SummaryWe are seeking an experienced Engineering Manager, GPU Infrastructure to lead the planning, deployment, integration, and operational readiness of large-scale AI infrastructure environments. This role is responsible for delivering production-grade GPU clusters that support AI training, inference, and high-performance computing workloads across cloud, hybrid, and on-premises environments.The ideal candidate brings deep technical expertise in GPU infrastructure, networking, storage, automation, and datacenter deployment, combined with strong program leadership and cross-functional execution skills. This leader will oversee end-to-end AI cluster deployment initiatives, including hardware integration, rack-and-stack operations, provisioning automation, performance validation, and operational handoff.The role requires hands-on familiarity with modern AI infrastructure tooling and architectures, including Canonical MaaS, VAST Data storage platforms, and both InfiniBand and Ethernet-based GPU networking fabrics.Key ResponsibilitiesAI and GPU Cluster Deployment & DeliveryOversee and partake in deployment and integration of GPU-based compute platforms from NVIDIA and other accelerator vendorsLead and participate in end-to-end logical deployment of large-scale AI and GPU clusters in state of the art datacenters.Manage deployment programs spanning compute, storage, networking, power, cooling, and automation layers.Participate in cluster architecture review for AI training, inference and distributed compute workloadsCoordinate rack-and-stack and cabling sequencing, network deployment, burn-in testing, and cluster validation.Validate deployment readiness, topology consistency, GPU fabric performance, acceptance testing, and operational turnover processes.Establish repeatable and documented deployment methodologies and scalable operational standards.Networking & Fabric ManagementLead deployment and operational validation of high-performance GPU interconnects using InfiniBand and Ethernet GPU fabric architecturesEnsure proper implementation of: spile-leaf architectures, RDMA, network telemetry and performance tuningCoordinate closely with network engineering teams on topology implementation and performance optimization.Storage & Data InfrastructureCoordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads.Ensure successful implementation and operational optimization of data storage platformsValidate storage throughput, latency, and GPU data delivery performance.Automation & ProvisioningLead infrastructure automation initiatives for cluster provisioning and lifecycle management.Manage deployment tooling and orchestration platforms including:Infrastructure-as-Code frameworksAutomated imaging and provisioning systems (e.g. Canonical MaaS)Cluster monitoring and observability toolsDrive standardization and deployment automation to improve speed, reliability, and repeatability.Leadership & Program ManagementBuild and lead high-performing technical deployment and infrastructure engineering teams.Partner with datacenter operations, hardware vendors, networking teams, and AI platform engineering groups.Establish strong Project Management Office (PMO) partnership while driving consistent, accurate project updates across the team and systems (e.g. Jira)Develop operational procedures, documentation, and deployment best practices.Mentor engineers and technical leads across infrastructure domains.Qualifications RequiredBachelor's degree in Computer Science, Engineering, Information Technology, or related field (or equivalent experience).10+ years of infrastructure engineering or datacenter deployment experience.5+ years leading deployment or operations teams supporting large-scale AI, HPC, or GPU infrastructure.Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments.Strong expertise with:Canonical MaaSData storage platformsInfiniBand and Ethernet GPU fabricsNetwork architectureLinux systems administrationGPU server architecturesStrong understanding of:RDMA and RoCE networkingHigh-performance storage architecturesCluster automation and provisioningDatacenter infrastructure operationsProven ability to manage complex cross-functional infrastructure deployment programs.Preferred QualificationsExperience deploying NVIDIA DGX SuperPOD or similar AI infrastructure solutions.Familiarity with:NVIDIA networking technologiesSpectrum-X or Quantum platformsAI model training infrastructureLiquid cooling environmentsDCIM and observability platformsExperience in hyperscale, cloud, or AI infrastructure environments.Certifications in networking, Linux, Kubernetes, or cloud infrastructure are a plus.Key CompetenciesTechnical leadershipInfrastructure architectureProgram executionCross-functional collaborationVendor and stakeholder managementProblem-solving under operational pressureProcess improvement and automationExcellent communication and documentation skills5C Data Centers is an equal opportunity employer. We evaluate all qualified applicants without regard to race, religion, gender, age, national origin, disability, sexual orientation, veteran status, or other protected status.