{"schemaVersion":"jobsearcher.job.v1","id":"d8f639320ae5a4f4eb679cfb","url":"https://jobsearcher.com/jobs/d8f639320ae5a4f4eb679cfb","canonicalUrl":"https://jobsearcher.com/jobs/d8f639320ae5a4f4eb679cfb","title":"Cluster Engineer","description":"Position Summary\nWe are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.\n\nResponsibilities\n\nDesign, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.\n\nTune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.\n\nOptimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.\n\nBuild and support production AI infrastructure running hundreds to thousands of GPUs.\n\nAnalyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.\n\nPerform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.\n\nDesign and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.\n\nConfigure and tune distributed AI software stacks including:\n\nPyTorch\n\nNCCL\n\nCUDA\n\nUCX\n\nMPI\n\nSlurm\n\nPyxis/Enroot\n\nOptimize GPU scheduling and resource allocation for both training and inference environments.\n\nDevelop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.\n\nIdentify performance regressions and troubleshoot distributed training issues at scale.\n\nOptimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.\n\nWork closely with ML engineers to improve training scalability and inference efficiency.\n\nCreate automation to deploy, validate, benchmark, and monitor GPU clusters.\n\nEvaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.\n\nRequired Qualifications\n\n7+ years designing or operating large-scale Linux infrastructure.\n\n5+ years supporting production GPU clusters for AI or HPC workloads.\n\nDemonstrated experience building multi-node GPU training environments from the ground up.\n\nDeep expertise with distributed PyTorch training.\n\nExtensive experience troubleshooting and optimizing NCCL communications.\n\nStrong understanding of distributed AI communication patterns, including:\n\nAllReduce\n\nReduceScatter\n\nAllGather\n\nBroadcast\n\nPoint-to-point communications\n\nExperience benchmarking distributed training using tools such as:\n\nnccl-tests\n\nNVIDIA DCGM\n\nNsight Systems\n\nMLPerf (preferred)\n\nStrong understanding of GPU memory management, including:\n\nKV Cache\n\nActivation checkpointing\n\nTensor Parallelism\n\nPipeline Parallelism\n\nData Parallelism\n\nExperience optimizing LLM inference throughput, including:\n\nTokens/sec optimization\n\nBatch sizing\n\nContinuous batching\n\nKV cache tuning\n\nMemory bandwidth optimization\n\nExperience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.\n\nExpert-level Linux systems administration skills.\n\nExperience with Slurm workload manager.\n\nExperience using Pyxis and Enroot for containerized GPU workloads.\n\nStrong scripting skills using Python and Bash.\n\nTechnical Expertise\nAI Frameworks\n\nPyTorch\n\nCUDA\n\nNCCL\n\nTriton (preferred)\n\nTensorRT-LLM (preferred)\n\nCluster Scheduling\n\nSlurm\n\nPyxis\n\nEnroot\n\nGPU Networking\n\nInfiniBand\n\nRoCE v2\n\nRDMA\n\nGPUDirect RDMA\n\nGPUDirect Storage\n\nUCX\n\nMPI\n\nNetwork topology optimization\n\nCongestion control\n\nQoS\n\nECN/PFC\n\nHigh-speed Ethernet (200/400/800 GbE)\n\nStorage\n\nParallel file systems\n\nDistributed storage\n\nObject storage\n\nNVMe\n\nCheckpoint optimization\n\nDataset staging\n\nGPUDirect Storage\n\nStorage bandwidth optimization\n\nMetadata performance\n\nPerformance Engineering\n\nNCCL benchmarking\n\nMulti-node scaling analysis\n\nGPU utilization optimization\n\nCommunication/computation overlap\n\nNUMA optimization\n\nCPU affinity\n\nPCIe topology\n\nGPU topology (NVLink/NVSwitch)\n\nMemory bandwidth analysis\n\nEnd-to-end performance profiling\n\nPreferred Qualifications\n\nExperience deploying AI workloads on Kubernetes.\n\nExperience with NVIDIA GPU Operator.\n\nExperience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).\n\nExperience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.\n\nExperience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.\n\nFamiliarity with MLPerf benchmarking.\n\nExperience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.\n\nExperience automating infrastructure using Ansible, Terraform, or similar tools.\n\nExperience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.\n\n#J-18808-Ljbffr","company":"Stn","rawCompany":"stn","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-18T03:37:46.786Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Cluster Engineer","description":"Position Summary\nWe are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.\n\nResponsibilities\n\nDesign, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.\n\nTune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.\n\nOptimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.\n\nBuild and support production AI infrastructure running hundreds to thousands of GPUs.\n\nAnalyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.\n\nPerform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.\n\nDesign and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.\n\nConfigure and tune distributed AI software stacks including:\n\nPyTorch\n\nNCCL\n\nCUDA\n\nUCX\n\nMPI\n\nSlurm\n\nPyxis/Enroot\n\nOptimize GPU scheduling and resource allocation for both training and inference environments.\n\nDevelop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.\n\nIdentify performance regressions and troubleshoot distributed training issues at scale.\n\nOptimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.\n\nWork closely with ML engineers to improve training scalability and inference efficiency.\n\nCreate automation to deploy, validate, benchmark, and monitor GPU clusters.\n\nEvaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.\n\nRequired Qualifications\n\n7+ years designing or operating large-scale Linux infrastructure.\n\n5+ years supporting production GPU clusters for AI or HPC workloads.\n\nDemonstrated experience building multi-node GPU training environments from the ground up.\n\nDeep expertise with distributed PyTorch training.\n\nExtensive experience troubleshooting and optimizing NCCL communications.\n\nStrong understanding of distributed AI communication patterns, including:\n\nAllReduce\n\nReduceScatter\n\nAllGather\n\nBroadcast\n\nPoint-to-point communications\n\nExperience benchmarking distributed training using tools such as:\n\nnccl-tests\n\nNVIDIA DCGM\n\nNsight Systems\n\nMLPerf (preferred)\n\nStrong understanding of GPU memory management, including:\n\nKV Cache\n\nActivation checkpointing\n\nTensor Parallelism\n\nPipeline Parallelism\n\nData Parallelism\n\nExperience optimizing LLM inference throughput, including:\n\nTokens/sec optimization\n\nBatch sizing\n\nContinuous batching\n\nKV cache tuning\n\nMemory bandwidth optimization\n\nExperience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.\n\nExpert-level Linux systems administration skills.\n\nExperience with Slurm workload manager.\n\nExperience using Pyxis and Enroot for containerized GPU workloads.\n\nStrong scripting skills using Python and Bash.\n\nTechnical Expertise\nAI Frameworks\n\nPyTorch\n\nCUDA\n\nNCCL\n\nTriton (preferred)\n\nTensorRT-LLM (preferred)\n\nCluster Scheduling\n\nSlurm\n\nPyxis\n\nEnroot\n\nGPU Networking\n\nInfiniBand\n\nRoCE v2\n\nRDMA\n\nGPUDirect RDMA\n\nGPUDirect Storage\n\nUCX\n\nMPI\n\nNetwork topology optimization\n\nCongestion control\n\nQoS\n\nECN/PFC\n\nHigh-speed Ethernet (200/400/800 GbE)\n\nStorage\n\nParallel file systems\n\nDistributed storage\n\nObject storage\n\nNVMe\n\nCheckpoint optimization\n\nDataset staging\n\nGPUDirect Storage\n\nStorage bandwidth optimization\n\nMetadata performance\n\nPerformance Engineering\n\nNCCL benchmarking\n\nMulti-node scaling analysis\n\nGPU utilization optimization\n\nCommunication/computation overlap\n\nNUMA optimization\n\nCPU affinity\n\nPCIe topology\n\nGPU topology (NVLink/NVSwitch)\n\nMemory bandwidth analysis\n\nEnd-to-end performance profiling\n\nPreferred Qualifications\n\nExperience deploying AI workloads on Kubernetes.\n\nExperience with NVIDIA GPU Operator.\n\nExperience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).\n\nExperience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.\n\nExperience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.\n\nFamiliarity with MLPerf benchmarking.\n\nExperience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.\n\nExperience automating infrastructure using Ansible, Terraform, or similar tools.\n\nExperience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.\n\n#J-18808-Ljbffr","datePosted":"2026-08-18T03:37:46.786Z","dateModified":"2026-08-18T03:37:46.786Z","hiringOrganization":{"@type":"Organization","name":"Stn","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d8f639320ae5a4f4eb679cfb"},"url":"https://jobsearcher.com/jobs/d8f639320ae5a4f4eb679cfb"}}