GPU Kernel Engineer
About the CompanyWe're partnering with a well-funded frontier AI research lab that is building the next generation of foundation models for reliable, real-world AI systems. Rather than focusing solely on benchmark performance, the team is developing AI capable of robust reasoning, decision-making, and autonomous execution at scale. Founded by researchers and engineers from some of the world's leading AI organizations, the company is backed by top-tier investors and is tackling some of the hardest systems challenges in modern machine learning.About the RoleThey're looking for exceptional GPU Kernel Engineers, particularly PhD graduates or researchers with deep expertise in GPU computing, high-performance systems, and large-scale machine learning. You'll work at the lowest levels of the AI stack, designing and optimizing custom GPU kernels that accelerate the training and inference of frontier-scale language models.ResponsibilitiesDesign, implement, and optimize high-performance CUDA kernels for training and inferenceProfile end-to-end model execution and eliminate performance bottlenecksImprove GPU utilization, memory efficiency, throughput, and latencyWork closely with AI researchers to optimize new model architecturesPush modern GPU hardware to its limits across large-scale distributed training systemsQualificationsThis role is particularly suited to PhD graduates from leading universities with research in areas such as:GPU ComputingHigh Performance Computing (HPC)Parallel ComputingComputer ArchitectureMachine Learning SystemsDistributed SystemsProgramming Languages & CompilersScientific ComputingRequired SkillsPhD from a top university (MIT, Stanford, Berkeley, CMU, UIUC, Princeton, Georgia Tech, University of Washington, ETH Zurich, Oxford, Cambridge, Toronto, etc.)Deep expertise in CUDA and GPU programmingStrong understanding of GPU architecture, memory hierarchy, scheduling, and parallelismExperience optimizing deep learning training or inference workloadsHands-on experience with large-scale LLM training (beyond research projects or hobby work)Excellent C++ and CUDA development skillsAbility to reason from first principles about performance and systems optimizationPreferred SkillsTriton, CUTLASS, CuTe DSL, or PyTorch internalsDistributed training frameworks (DeepSpeed, FSDP, Megatron, NCCL)Compiler technologies (LLVM, MLIR, Triton, XLA)CUDA GraphsQuantization or inference optimizationExperience using Nsight Systems or Nsight Compute for performance analysisThis is an excellent opportunity for outstanding PhD graduates who want to apply their research in GPU systems, HPC, ML systems, or computer architecture to cutting-edge frontier AI.