Technical Program Manager – AI Infrastructure / GPU Clusters
Technical Program Manager – AI Infrastructure / GPU ClustersLocation: APAC / USCompany: GMI CloudAbout GMI CloudGMI Cloud is building next-generation AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions.We are looking for a Technical Program Manager (TPM) to drive the deployment and delivery of GPU cluster infrastructure. This role will work at the intersection of AI hardware platforms, high-performance networking, and data center infrastructure, coordinating across solution architects, engineering teams, vendors, and contractors to deliver production-ready AI clusters.ResponsibilitiesGPU Cluster DeploymentLead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.Infrastructure Architecture CollaborationWork closely with Infrastructure Solution Architects (SA) to define:GPU server platform selectionNetwork architecture for distributed GPU clustersStorage integration and cluster infrastructure designSupport development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.System IntegrationDrive system integration for large-scale GPU clusters, including:Rack elevation planningGPU server deployment and configurationHigh-speed network topology implementationPower and cooling readinessEnsure deployments align with vendor reference architectures and validated cluster designs.Contractor & Field Deployment ManagementWork closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.Coordinate and oversee field deployment activities such as:Structured cabling installationRack installation and equipment mountingNetwork and power connectivity preparationHardware staging and deployment logisticsCluster Validation & Performance TestingCoordinate cluster bring-up and validation activities including:Single-node GPU validationMulti-node cluster deploymentGPU interconnect validation (P2P, RDMA)Drive cluster benchmarking, stress testing, and performance verification before production release.Operational ReadinessEnsure deployed GPU clusters are fully ready for production workloads by driving:Hardware and network validationMonitoring and telemetry integrationOperational documentation and runbooksHandover to operations teamsRequired Qualifications5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure deliveryExperience with GPU cluster deployments or high-performance computing environmentsFamiliarity with GPU server architecture and distributed computing infrastructureExperience working with Infrastructure Solution Architects to define system architecture and hardware BOMExperience managing data center hardware deployments and system integrationAbility to coordinate multi-vendor infrastructure projects across regionsPreferred QualificationsExperience deploying large-scale AI infrastructure or GPU clustersFamiliarity with:InfiniBand / RoCE / high-speed Ethernet networkingGPU interconnect validation (P2P / RDMA)rack elevation and high-density rack deploymentExperience with cluster validation and performance benchmarkingBackground as Systems Engineer, HPC Engineer, or Infrastructure ArchitectExperience working in AI infrastructure, cloud infrastructure, or hyperscale data centersNice to HaveExperience deploying liquid-cooled GPU clusters or high-power racksExperience working with NVIDIA AI infrastructure platformsFamiliarity with AI training environments and distributed workloads