Technical Program Manager - Performance & Benchmarking
Overview
You will lead cross-functional programs to ensure CoreWeave’s AI/ML platform runs at peak performance and reliability. Collaborating with engineering, infrastructure, product, and GTM teams, you’ll drive benchmarking, validation, and launch readiness across hardware, clusters, and models. You’ll turn performance insights into roadmap priorities and customer-ready capabilities, shaping how researchers and enterprises use CoreWeave. This role offers impact across infrastructure readiness and measurable performance improvements.
Compensation / BenefitsMedical, dental, and vision insurance - 100% paidCompany-paid Life InsuranceDisability insurance (short and long-term)Flexible Spending Account (FSA) / Health Savings Account (HSA)Tuition ReimbursementEmployee Stock Purchase Program (ESPP)
ResponsibilitiesDrive end-to-end program execution for performance and benchmarking initiatives including validation, testing, benchmarking, observability, and launch readinessPartner with engineering and infrastructure to verify new hardware platforms, clusters, and software environments meet performance and stability standardsLead cross-functional efforts to operationalize benchmarking frameworks measuring model performance, runtime efficiency, GPU utilization, and reliabilityCoordinate dependencies across platform engineering, infrastructure, capacity, product, and GTM to translate findings into roadmap priorities and customer readinessBuild mechanisms for release readiness, benchmark planning, risk management, escalation, and post-launch reviews for performance-focused initiativesEstablish dashboards, cadences, and success metrics to improve performance visibility and validation coverageHelp prioritize performance bottlenecks, test gaps, and benchmark requests by aligning stakeholders on goals and tradeoffsCreate clarity across ambiguous programs by aligning teams around performance goals, validation criteria, and milestones
Key requirementsBachelor’s degree in CS/Engineering or equivalent experience5+ years of technical program management in cloud infrastructure, distributed systems, HPC, or AI/ML platformsExperience leading large cross-functional programs including performance engineering, benchmarking, or infrastructure readinessStrong technical fluency in distributed systems, GPU/accelerator-based infrastructure, workload performance measurement, and large-scale operationsAbility to define program metrics and drive outcomes in performance, reliability, scale, or operational maturityExcellent communication with cross-functional influence over engineering, product, and infrastructureExperience with AI/ML benchmarking, performance analysis, or infrastructure validation for training and inference workloadsFamiliarity with GPU cluster architecture, observability, hardware bring-up, and bottleneck analysisUnderstanding of benchmarking methodologies, reproducibility, test coverage, and tradeoffs between performance, stability, utilization, and readinessExperience building launch processes, release governance, dependency management, and operational reviews in fast-scaling environmentsAbility to translate technical performance data into actionable decisions for product/customer/go-to-market audiencesExcellent communicationCross-functional collaborationAnalytical mindsetDistributed systemsGPU/accelerator-based infrastructurePerformance measurement and benchmarking