Technical Program Manager – Infrastructure Engineering
About USGMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents. Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter. From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud. One cloud for compute, inference, and agents.About the RoleGMI Cloud is building next-generation AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions. We combine infrastructure automation, cloud-native orchestration, and high-performance hardware to make GPU compute simple, reliable, and cost-effective. As we scale, we’re hiring a Technical Program Manager to coordinate cross-functional engineering efforts and help deliver platform reliability, performance, and developer experience.As an Infra Engineering TPM you will be the connective tissue between SRE, Platform Developers, Architects, and the broader engineering organization. You will own planning, execution, and communication for infrastructure programs—ensuring work aligns to product priorities and is delivered on time with high quality. You’ll support the Head of Infrastructure to turn strategic priorities into actionable roadmaps and to unblock teams operating in a fast-moving start-up environment.Responsibilities1) Program leadershipLead end-to-end technical programs across infrastructure domains (GPU orchestration, cluster management, networking, storage, telemetry, provisioning pipelines).Align program scope, milestones, dependencies, success metrics, and delivery timelines.Drive program rituals: kickoff, sprint alignment, risk reviews, and postmortems.2) Cross-team coordination & stakeholder managementProactively support the Head of Infrastructure in setting priorities, sequencing work, meeting coordination, and driving execution across teams.Serve as the coordinator within infrastructure organization: SRE, Platform Developers, Architects.Proactively surface and manage dependencies, blockers, and risks across teams.Communicate program status, trade-offs, and timelines to the Head of Infrastructure and other stakeholders.3) Delivery and processManage roadmaps and release plans for multi-team infrastructure initiatives; work directly with the Head of Infrastructure to align execution with strategic goals and adjust priorities as needed.Implement scalable program management practices (OKRs, milestones, RACI, risk logs, SLAs).Ensure engineering deliverables meet performance, reliability, security, and cost targets.4) Technical grounding & decision supportMaintain a strong technical understanding of GPU infrastructure, ensure architecture and program decisions align with NVIDIA NCP (NVIDIA Cloud Platform) reference architecture and best practices where applicable — including hardware configurations, networking/topology, telemetry, and security recommendations.Facilitate architecture reviews and help translate architecture into executable work for platform engineers and SREs.5) Reliability & operations enablementCoordinate SRE-driven reliability initiatives (SLOs, incident response playbooks, SOP).Help prioritize reliability work vs. feature work and plan operational runbooks, testing, and rollout strategies.6) Metrics & continuous improvementDefine and track key KPIs (uptime, mean time to recovery, provisioning latency).Use data to drive prioritization and continuous improvement cycles.QualificationMinimum QualificationBachelor’s degree in Computer Science, Engineering, or equivalent experience.5+ years of technical program management experience at an infrastructure- or platform-oriented organization (start-up experience strongly preferred).Strong technical background with hands-on familiarity in at least one of: GPU compute stacks, Kubernetes, cluster orchestration, or cloud infrastructure.Demonstrated experience coordinating across SRE, platform engineering, and architecture teams to deliver complex multi-quarter programs.Proven track record of delivering projects with multiple engineering teams and external dependencies.Excellent written and verbal communication—able to summarize technical trade-offs and program status for executive stakeholders.Strong organizational skills: roadmaps, dependency mapping, risk management, and prioritization.Data-driven: experience defining, monitoring, and improving KPIs and SLAs.Comfortable with ambiguity and rapid change in a start-up environment.Experience with incident management, postmortems, and on-call processes.Nice-to-havesDirect experience with GPU orchestration tools (Device Plugin, MIG, NCCL, CUDA drivers), Kubernetes GPU scheduling frameworks, or custom scheduler implementations.Experience with on-prem/cloud deployments.Familiarity with infrastructure-as-code, CI/CD, and observability tooling (Prometheus, Grafana, OpenTelemetry).Prior experience at a GPU-focused company, ML infrastructure team, or HPC environment.