{"schemaVersion":"jobsearcher.job.v1","id":"dc9675dca94e28a5f4e95a91","url":"https://jobsearcher.com/jobs/dc9675dca94e28a5f4e95a91","canonicalUrl":"https://jobsearcher.com/jobs/dc9675dca94e28a5f4e95a91","title":"Senior CUDA Engineer - Equivariant ML & GPU Kernels","description":"NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Join our group and discover how you can develop a lasting impact on the world.\n\nNVIDIA BioNeMo is building the computational foundation for the next generation of biological discovery. We are looking for a Senior Software Engineer to join the cuEquivariance team - an NVIDIA library that accelerates geometric neural networks on NVIDIA GPUs, enabling researchers in molecular biology, materials science, and physics to train and deploy equivariant models at scale. This team builds and ships the production GPU kernels and software interfaces that power equivariant deep learning throughout the scientific field. The work spans CUDA kernel engineering, Python library development involving both PyTorch and JAX, and direct collaboration with research teams and external framework developers. If you want to work where GPU computing meets graph-based deep learning, this is the role for you. Your work will run in production pipelines across the scientific community.\n\nWhat You Will Be Doing:\nBuild, implement, and optimize CUDA kernels for equivariant neural network primitives - tensor products, segmented polynomials, and triangle-based operations - targeting peak performance across NVIDIA GPU generations.\nBe responsible for the end-to-end delivery of GPU-accelerated geometric ML primitives: from implementation to validated, production-quality software that external frameworks depend on.\nBuild and maintain the interfaces for PyTorch and JAX that expose cuEquivariance primitives to application developers and researchers.\nDrive CI/CD infrastructure for multi-GPU kernel builds, automated correctness testing, and performance regression tracking.\nCollaborate with Applied Science and research teams to evaluate new equivariant architectures and translate prototypes into production kernels.\nEngage directly with third-party framework developers and partners to align on interfaces and ensure delivered software integrates cleanly into production pipelines.\n\nWhat We Need to See:\n6+ years of software engineering experience with a strong background in CUDA and GPU programming.\nDeep proficiency in C++ and Python; experience building and shipping production libraries used by external developers.\nGood foundation in GPU computing: memory hierarchy, warp-level execution, occupancy, and performance profiling methodology.\nExperience building or chipping in to production scientific software libraries, ML frameworks, or developer-facing GPU APIs.\nFamiliarity with concepts in geometric machine learning - equivariance, group representations, irreducible representations, or tensor products - sufficient to work efficiently in the domain.\nBS/MS in Computer Science, Physics, Applied Mathematics, or a related field, or equivalent experience.\n\nWays to Stand Out from the Crowd:\nYou have chipped in to or deeply used a major neural network framework that respects equivariance: e3nn, MACE, NequIP, SE(3)-Transformers, or similar.\nHands-on experience with Triton kernel development or other GPU kernel authoring tools alongside CUDA.\nExperience with mixed-precision or tensor-core-aware algorithm design for scientific or ML workloads.\nPhD or equivalent experience in computational chemistry, biophysics, physics, or computer science with a focus on geometric deep learning or HPC.\nContributions to open-source geometric ML or GPU computing projects.\n\nWidely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/\n\nYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.\n\nYou will also be eligible for equity and benefits .\n\nApplications for this job will be accepted at least until May 26, 2026.\n\nThis posting is for an existing vacancy.\n\nNVIDIA uses AI tools in its recruiting processes.\n\nNVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.\n#J-18808-Ljbffr","company":"NVIDIA","rawCompany":"nvidia","city":"Santa Clara","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-07-16T03:33:37.771Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior CUDA Engineer - Equivariant ML & GPU Kernels","description":"NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Join our group and discover how you can develop a lasting impact on the world.\n\nNVIDIA BioNeMo is building the computational foundation for the next generation of biological discovery. We are looking for a Senior Software Engineer to join the cuEquivariance team - an NVIDIA library that accelerates geometric neural networks on NVIDIA GPUs, enabling researchers in molecular biology, materials science, and physics to train and deploy equivariant models at scale. This team builds and ships the production GPU kernels and software interfaces that power equivariant deep learning throughout the scientific field. The work spans CUDA kernel engineering, Python library development involving both PyTorch and JAX, and direct collaboration with research teams and external framework developers. If you want to work where GPU computing meets graph-based deep learning, this is the role for you. Your work will run in production pipelines across the scientific community.\n\nWhat You Will Be Doing:\nBuild, implement, and optimize CUDA kernels for equivariant neural network primitives - tensor products, segmented polynomials, and triangle-based operations - targeting peak performance across NVIDIA GPU generations.\nBe responsible for the end-to-end delivery of GPU-accelerated geometric ML primitives: from implementation to validated, production-quality software that external frameworks depend on.\nBuild and maintain the interfaces for PyTorch and JAX that expose cuEquivariance primitives to application developers and researchers.\nDrive CI/CD infrastructure for multi-GPU kernel builds, automated correctness testing, and performance regression tracking.\nCollaborate with Applied Science and research teams to evaluate new equivariant architectures and translate prototypes into production kernels.\nEngage directly with third-party framework developers and partners to align on interfaces and ensure delivered software integrates cleanly into production pipelines.\n\nWhat We Need to See:\n6+ years of software engineering experience with a strong background in CUDA and GPU programming.\nDeep proficiency in C++ and Python; experience building and shipping production libraries used by external developers.\nGood foundation in GPU computing: memory hierarchy, warp-level execution, occupancy, and performance profiling methodology.\nExperience building or chipping in to production scientific software libraries, ML frameworks, or developer-facing GPU APIs.\nFamiliarity with concepts in geometric machine learning - equivariance, group representations, irreducible representations, or tensor products - sufficient to work efficiently in the domain.\nBS/MS in Computer Science, Physics, Applied Mathematics, or a related field, or equivalent experience.\n\nWays to Stand Out from the Crowd:\nYou have chipped in to or deeply used a major neural network framework that respects equivariance: e3nn, MACE, NequIP, SE(3)-Transformers, or similar.\nHands-on experience with Triton kernel development or other GPU kernel authoring tools alongside CUDA.\nExperience with mixed-precision or tensor-core-aware algorithm design for scientific or ML workloads.\nPhD or equivalent experience in computational chemistry, biophysics, physics, or computer science with a focus on geometric deep learning or HPC.\nContributions to open-source geometric ML or GPU computing projects.\n\nWidely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/\n\nYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.\n\nYou will also be eligible for equity and benefits .\n\nApplications for this job will be accepted at least until May 26, 2026.\n\nThis posting is for an existing vacancy.\n\nNVIDIA uses AI tools in its recruiting processes.\n\nNVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:33:37.771Z","dateModified":"2026-07-16T03:33:37.771Z","hiringOrganization":{"@type":"Organization","name":"NVIDIA","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Santa Clara","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"dc9675dca94e28a5f4e95a91"},"url":"https://jobsearcher.com/jobs/dc9675dca94e28a5f4e95a91"}}