Senior Systems Software Engineer - GPU Performance at Scale
Overview
As a Senior Performance Engineer at NVIDIA, you will shape at-scale AI system performance and datacenter applications. You will partner across HW, FW, and software teams to design and tune scalable AI workloads, delivering tools and insights that accelerate datacenter performance. Your work spans collaborating with researchers, developers, and customers to craft optimized workflows and new solutions for large-scale AI systems. You’ll lead performance practices in multi-product environments and translate complex issues into actionable fixes.
ResponsibilitiesLead performance practices for large-scale infrastructure and deliver tools and flows to validate and improve datacenter productsAlign next-generation AI workloads with future datacenter designs through early HW/FW/SW engagementProvide continuous performance insights for AI workloads across evolving environments and identify improvements and regressionsDecompose complex performance or stability issues into minimal reproduction cases to reach root causeCollaborate with SW/FW teams (BMC/SBIOS/OS/drivers) to develop best-in-class practices for AI workloads at scaleEngage with engineering and research communities supporting HPC or DL to resolve firmware/software issues for AI performance
Key requirements8+ years of experience in accelerated computing for datacenter/container computing solutionsProven understanding of accelerated computing software stacks and DL frameworks (CUDA, PyTorch)Experience with modern cloud and container-based Enterprise architecturesProficiency in C/C++/Python/Bash scriptingExperience with CPU architectureExperience with container tech and Linux-based OSesUnderstanding of collective communication patterns in AI workloadsExperience collaborating with engineering or academic research communities in HPC or DLStrong verbal and written communicationTeamwork and collaborationAnalytical mindset and action-orientedCUDAPyTorchcloud and container-based architectures