JOBSEARCHER

Senior Deep Learning Software Engineer, Inference

NVIDIAOakland, CAL6 LeadSeptember 15th, 2026
Overview In this role you will design, implement and optimize GPU-accelerated software powering state-of-the-art AI applications. You contribute to high-performance DL inference frameworks (SGLang, vLLM) and related tools, shaping reliable model serving at scale. You work across NVIDIA accelerators and open-source components to drive efficient deployment of cutting-edge language models. You will collaborate with cross-functional teams to apply the latest algorithms and achieve performance breakthroughs in inference. This is a mission-driven opportunity to impact large-scale AI systems and enterprise deployments. Compensation / Benefitsequitybenefitsremote work option ResponsibilitiesPerformance optimization, analysis, and tuning of DL models across domains (LLM, Multimodal, Generative AI)Scale DL model performance across architectures and NVIDIA acceleratorsContribute features and code to NVIDIA inference libraries (vLLM, SGLang, FlashInfer, LLM solutions)Collaborate with cross-functional teams across frameworks and NVIDIA libraries to innovate inference solutions Key requirements5+ years of relevant software development experienceExcellent C/C++ programming and software design skillsPython experience is a plusExperience with training, deploying or optimizing DL model inference in production is a plusPerformance modeling, profiling, debugging, and CPU/GPU architectural knowledge is a plusMasters or PhD (or equivalent experience) in Computer Engineering, Computer Science, EECS or AICollaborative and cross-functional teamworkStrong debugging and profiling mindsetResult-oriented with a focus on performance impactCUDACUTLASSNCCL