Lead GPU Cluster Solution Architect
ABOUT THE ROLEAxe Compute is seeking a Lead Architect, GPU Cluster Solutions to design GPU cluster configurations for prospective and signed engagements, translating client requirements and NVIDIA Reference Architecture into a buildable, supportable design, spanning compute, storage, networking, software and spares strategy for support.ROLE AT A GLANCEMandate: Own end-to-end technical design of GPU cluster deployments, from client requirements through NVIDIA Reference Architecture compliance and sparing strategy.Scope: Cluster design, network architecture (InfiniBand/RoCE/Ethernet), sparing and spares planning, connectivity design (internet/VPN/firewall, dedicated circuits), design adjustments for site and hardware constraints.Key Outcomes: Designs that meet client SLAs and NVIDIA Reference Architecture standards, sparing plans that protect uptime commitments, designs that account for real-world site and hardware lead-time constraints.WHAT YOU WILL OWNCluster Design & Reference ArchitectureDesign GPU cluster configurations (compute, storage, networking) against NVIDIA Reference Architecture for each signed engagement.Translate client technical requirements into a complete bill of design, including all necessary compute, storage, and networking components.Network ArchitectureDesign network topology and fabric selection, including InfiniBand, RoCE, and Ethernet options, appropriate to each client's workload and performance requirements.Incorporate internet, VPN, and firewall connectivity requirements into cluster designs.Design dedicated point-to-point network requirements where needed, including protected optical circuits and similar dedicated connectivity.Sparing & Availability StrategyFormulate and own the hot/cold sparing plan for each deployment to meet contracted SLA commitments.Adjust sparing and design assumptions based on data center power/cooling parameters and hardware lead-time constraints.Design Adaptation & Site ConstraintsAdjust cluster designs to fit site-specific power, cooling, and space constraints identified by the Data Center Procurement and Operations Director.Work with Supply Chain to align design decisions with realistic hardware delivery timing.Cross-Functional CollaborationPartner with the VP, Deployments and Deployment Program Manager to ensure designs translate cleanly into buildable, trackable project plans.Support acceptance test design and criteria definition, ensuring test procedures validate the as-designed architecture.REQUIRED QUALIFICATIONS7+ years in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure.Deep working knowledge of NVIDIA Reference Architecture (HGX, NVL72) and GPU cluster design principles.Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet fabric design.Experience designing sparing/spares strategies for mission-critical infrastructure.Experience incorporating firewall, VPN, and dedicated circuit (e.g., protected optical) requirements into network designs.Experience with high-speed shared storage solutions (e.g., Weka, Vast, DDN).PREFERRED QUALIFICATIONSExperience designing clusters for large enterprise clients or neoclouds, not just internal infrastructure.Familiarity with NVIDIA NCP program requirements and certification processes.Experience with capacity or sparing modeling tools.