{"schemaVersion":"jobsearcher.job.v1","id":"d152b2c00f319c6bd1af704c","url":"https://jobsearcher.com/jobs/d152b2c00f319c6bd1af704c","canonicalUrl":"https://jobsearcher.com/jobs/d152b2c00f319c6bd1af704c","title":"Senior Inference Engineer, GPU Kernel Optimization","description":"Overview\nIn this role you push the performance frontier of LLM inference by optimizing GPU kernels and building end-to-end performance tooling. You will work across kernel, compiler, hardware, and framework teams to surface bottlenecks and ship measurable gains. You’ll apply AI-driven analysis to diagnose gaps and validate improvements with silicon measurements. This role combines hands-on profiling, performance modeling, and agentic optimization to scale NVIDIA’s LLM inference stack.\n\nCompensation / Benefitsequitycompetitive base salarycomprehensive benefits packagecareer growth and impactflexible work arrangementsstructured bonus or incentive programs\nResponsibilitiesDevelop and drive GPU kernel microbenchmarking with real-silicon fidelityLead end-to-end model performance analysis to link evidence to production throughput and latencyApply AI-driven analysis for kernel optimization and validate findings with silicon measurements\nKey requirementsMaster's in Computer Science, Computer Engineering, or related field (or equivalent experience)6+ years of relevant industry experienceExperience building or directing agentic AI systems (code generation, automated optimization, or multi-step reasoning workflows)Strong Python and C++ programming skills with solid software engineering fundamentalsHands-on GPU profiling with CUPTI, NSYS, and NCU; ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime schedulingDirect experience with LLM inference frameworks (TRT-LLM, SGLang, or vLLM) and understanding of kernel choice impact on throughput/latencyKnowledge of GPU kernel optimization (CUDA, CUTLASS, Triton) and ability to read PTX or SASSCollaborative mindsetStrong communication of complex technical conceptsProblem-solving mindsetGPU kernel optimizationCUDA/CUTLASS/TritonPTX/SASS reading","company":"NVIDIA","rawCompany":"nvidia","city":"Brooklyn","state":"NY","isRemote":false,"isActive":false,"createdAt":"2026-09-15T04:03:24.375Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"334413","title":"Semiconductor and Related Device Manufacturing","slug":"semiconductor-and-related-device-manufacturing"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Inference Engineer, GPU Kernel Optimization","description":"Overview\nIn this role you push the performance frontier of LLM inference by optimizing GPU kernels and building end-to-end performance tooling. You will work across kernel, compiler, hardware, and framework teams to surface bottlenecks and ship measurable gains. You’ll apply AI-driven analysis to diagnose gaps and validate improvements with silicon measurements. This role combines hands-on profiling, performance modeling, and agentic optimization to scale NVIDIA’s LLM inference stack.\n\nCompensation / Benefitsequitycompetitive base salarycomprehensive benefits packagecareer growth and impactflexible work arrangementsstructured bonus or incentive programs\nResponsibilitiesDevelop and drive GPU kernel microbenchmarking with real-silicon fidelityLead end-to-end model performance analysis to link evidence to production throughput and latencyApply AI-driven analysis for kernel optimization and validate findings with silicon measurements\nKey requirementsMaster's in Computer Science, Computer Engineering, or related field (or equivalent experience)6+ years of relevant industry experienceExperience building or directing agentic AI systems (code generation, automated optimization, or multi-step reasoning workflows)Strong Python and C++ programming skills with solid software engineering fundamentalsHands-on GPU profiling with CUPTI, NSYS, and NCU; ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime schedulingDirect experience with LLM inference frameworks (TRT-LLM, SGLang, or vLLM) and understanding of kernel choice impact on throughput/latencyKnowledge of GPU kernel optimization (CUDA, CUTLASS, Triton) and ability to read PTX or SASSCollaborative mindsetStrong communication of complex technical conceptsProblem-solving mindsetGPU kernel optimizationCUDA/CUTLASS/TritonPTX/SASS reading","datePosted":"2026-09-15T04:03:24.375Z","dateModified":"2026-09-15T04:03:24.375Z","hiringOrganization":{"@type":"Organization","name":"NVIDIA","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Brooklyn","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d152b2c00f319c6bd1af704c"},"url":"https://jobsearcher.com/jobs/d152b2c00f319c6bd1af704c"}}