JOBSEARCHER

Senior Inference Engineer, GPU Kernel Optimization

NVIDIABrooklyn, NYL6 LeadSeptember 15th, 2026
Overview In this role you push the performance frontier of LLM inference by optimizing GPU kernels and building end-to-end performance tooling. You will work across kernel, compiler, hardware, and framework teams to surface bottlenecks and ship measurable gains. You’ll apply AI-driven analysis to diagnose gaps and validate improvements with silicon measurements. This role combines hands-on profiling, performance modeling, and agentic optimization to scale NVIDIA’s LLM inference stack. Compensation / Benefitsequitycompetitive base salarycomprehensive benefits packagecareer growth and impactflexible work arrangementsstructured bonus or incentive programs ResponsibilitiesDevelop and drive GPU kernel microbenchmarking with real-silicon fidelityLead end-to-end model performance analysis to link evidence to production throughput and latencyApply AI-driven analysis for kernel optimization and validate findings with silicon measurements Key requirementsMaster's in Computer Science, Computer Engineering, or related field (or equivalent experience)6+ years of relevant industry experienceExperience building or directing agentic AI systems (code generation, automated optimization, or multi-step reasoning workflows)Strong Python and C++ programming skills with solid software engineering fundamentalsHands-on GPU profiling with CUPTI, NSYS, and NCU; ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime schedulingDirect experience with LLM inference frameworks (TRT-LLM, SGLang, or vLLM) and understanding of kernel choice impact on throughput/latencyKnowledge of GPU kernel optimization (CUDA, CUTLASS, Triton) and ability to read PTX or SASSCollaborative mindsetStrong communication of complex technical conceptsProblem-solving mindsetGPU kernel optimizationCUDA/CUTLASS/TritonPTX/SASS reading