Senior Deep Learning Inference Performance Architect
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
Senior Deep Learning Inference Performance ArchitectWe are looking for a Senior Performance Architect—a creative engineer who loves to squeeze every cycle of performance from deep learning software. The Inference Architecture team does groundbreaking hardware-software co-design work that focuses on accelerating AI inference workloads. In this role you will write performance-optimized low level code on today's GPUs, evaluate and improve state-of-the-art performance techniques in production large language model deployments, and help guide our future GPU architecture decisions. If you enjoy digging deep into GPU architecture details, are passionate about AI, and know where every cycle goes when you write highly tuned software, this role may be a great fit for you.What You'll Be DoingDevelop innovative GPU and system architectures to extend the state of the art in AI inference performance and efficiencyModel, analyze and prototype key deep learning algorithms and applicationsUnderstand and analyze the interplay of hardware and software architectures on future algorithms and applicationsWrite efficient software for AI inference, including CUDA kernels, framework-level code, and application-level codeCollaborate across the company to guide the direction of AI, working with software, research and product teamsWhat We Need To SeeA MS or PhD in a relevant discipline (CS, EE, Math) or equivalent experience, with 5+ years of relevant experienceStrong mathematical foundation in machine learning and deep learningExpert programming skills in C, C++, and PythonFamiliarity with GPU computing (CUDA or similar) and HPC (MPI, OpenMP)Strong knowledge and coursework in computer architectureWays To Stand Out From The CrowdBackground with systems-level performance modeling, profiling, and analysisExperience in characterizing and modeling system-level performance, executing comparison studies, and documenting and publishing resultsExperience in optimizing AI inference workloads with CUDA kernel developmentYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD – 287,500 USD for Level 4, and 224,000 USD – 356,500 USD for Level 5. You will also be eligible for equity and benefits.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.J-18808-Ljbffr