Senior Inference Engineer, GPU Kernel Optimization
Overview
In this role you push the performance frontier of LLM inference by optimizing GPU kernels and building end-to-end performance tooling. You will work across kernel, compiler, hardware, and framework teams to surface bottlenecks and ship measurable gains. You’ll apply AI-driven analysis to diagnose gaps and validate improvements with silicon measurements. This role combines hands-on profiling, performance modeling, and agentic optimization to scale NVIDIA’s LLM inference stack.
Compensation / Benefitsequitycompetitive base salarycomprehensive benefits packagecareer growth and impactflexible work arrangementsstructured bonus or incentive programs
ResponsibilitiesDevelop and drive GPU kernel microbenchmarking with real-silicon fidelityLead end-to-end model performance analysis to link evidence to production throughput and latencyApply AI-driven analysis for kernel optimization and validate findings with silicon measurements
Key requirementsMaster's in Computer Science, Computer Engineering, or related field (or equivalent experience)6+ years of relevant industry experienceExperience building or directing agentic AI systems (code generation, automated optimization, or multi-step reasoning workflows)Strong Python and C++ programming skills with solid software engineering fundamentalsHands-on GPU profiling with CUPTI, NSYS, and NCU; ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime schedulingDirect experience with LLM inference frameworks (TRT-LLM, SGLang, or vLLM) and understanding of kernel choice impact on throughput/latencyKnowledge of GPU kernel optimization (CUDA, CUTLASS, Triton) and ability to read PTX or SASSCollaborative mindsetStrong communication of complex technical conceptsProblem-solving mindsetGPU kernel optimizationCUDA/CUTLASS/TritonPTX/SASS reading