JOBSEARCHER

Sr. Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron

AmazonSan Jose, CAL6 LeadSeptember 15th, 2026
Overview In this role you will onboard and optimize state-of-the-art open-source and customer LLMs for inference on Trainium accelerators. You will drive performance, reliability, and usability across the Neuron inference stack while leading technical initiatives. Collaborating with customers and cross-functional teams, you’ll shape architecture, contribute to documentation, and mentor engineers to accelerate model deployment at scale. This is a hands-on, impact-focused opportunity to advance high-performance inference for a cutting-edge AI platform. Compensation / Benefitshealth insurance (medical, dental, vision)401(k) matchingpaid time offparental leaveRSUs and sign-on paymentscomprehensive benefits package ResponsibilitiesDeliver high-performance models using distributed inference librariesDrive performance optimization and system reliability across the Neuron ecosystemMentor team members and provide technical leadership across multiple work streamsDrive architectural decisions that impact the entire Neuron serving stackCollaborate with customers, product owners, and engineering teams to define technical strategyAuthor technical documentation, design proposals, and architectural guidelinesLead design reviews and architectural discussionsDebug complex performance issues across the stack with compiler and runtime teamsMentor junior engineers on system design and model optimization across model enablement teams Key requirements5+ years of programming in Java, C++, or C#, with OO design experience5+ years of leading design or architecture of new/existing systems5+ years of full software development lifecycle, including coding standards, code reviews, SCM, build, testing, operations5+ years of non-internship professional software development experienceExperience as a mentor, tech lead, or engineering team leaderteam leadershipmentorshipcross-functional collaborationproficiency in Java, C++, or C#distributed inference librariesperformance optimization and profiling