HPC Infrastructure Engineer
A seed-stage neocloud company is hiring an HPC Infrastructure Engineer to build and operate the foundation of its GPU and TPU cloud. You'll work across bare-metal systems, high-speed networking, cluster scheduling, storage, automation, and reliability, turning a large fleet of accelerators into a high-performance, dependable computing platform for customers. This is a hands-on role for someone who loves working across hardware and software, diagnosing tough performance problems, and building infrastructure from the ground up.What You'll DoDesign, deploy, and operate production TPU clustersProvision and manage Linux-based bare-metal servers at scaleBuild automated workflows for server installation, configuration, upgrades, and recoveryDeploy and operate cluster schedulers such as Slurm or KubernetesIntegrate high-performance networking using InfiniBand or RoCE/RDMABuild and operate high-throughput storage for distributed AI workloadsMonitor cluster health, accelerator utilization, network performance, storage performance, and job reliabilityDiagnose failures across GPUs, servers, firmware, networks, storage, schedulers, and customer workloadsImprove cluster utilization, training performance, fault tolerance, and recovery timeEstablish production practices for change management, incident response, capacity management, and operational readinessPartner with data-center, network, platform, security, and customer-facing teams to launch new capacityEvaluate and manage infrastructure vendors, hardware suppliers, and technical partnersParticipate in an on-call rotation as the production platform growsWhat We're Looking ForExperience building or operating production HPC, supercomputing, or large-scale bare-metal infrastructureStrong Linux systems-engineering and debugging skillsExperience automating physical server provisioning and configurationExperience with Slurm, Kubernetes, or another distributed workload schedulerUnderstanding of high-performance networking, distributed storage, and accelerator-based computingExperience with monitoring, incident response, performance analysis, and production reliabilityProficiency in Python, Go, Bash, or another infrastructure automation languageAbility to diagnose problems across layers rather than treating compute, networking, storage, and software as separate systemsComfort working directly with hardware, vendors, data-center teams, and customers in an early-stage environmentEspecially ValuableNVIDIA GPU clusters, TPU systems, CUDA, NCCL, or collective communicationsInfiniBand, RoCEv2, RDMA, GPUDirect, or high-performance EthernetSlurm administration and scheduling optimizationKubernetes for accelerator workloadsPXE, Redfish, IPMI, BMCs, and bare-metal lifecycle managementAnsible, Terraform, or other infrastructure-as-code systemsLustre, Spectrum Scale/GPFS, Ceph, BeeGFS, or other distributed storage systemsPrometheus, Grafana, OpenTelemetry, or infrastructure observability toolingGPU health monitoring, firmware management, and hardware failure diagnosisDistributed AI training and inference workloadsPerformance benchmarking and optimization across compute, network, and storageExperience at a neocloud, hyperscaler, AI lab, national laboratory, or HPC center