{"schemaVersion":"jobsearcher.job.v1","id":"095902ffc9739bd6e7ed40ab","url":"https://jobsearcher.com/jobs/095902ffc9739bd6e7ed40ab","canonicalUrl":"https://jobsearcher.com/jobs/095902ffc9739bd6e7ed40ab","title":"Sr. Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron","description":"Sr. Software Development EngineerWe develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.\r\nAs a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.\r\nKey job responsibilities:\r\nDeliver high-performance models using distributed inference libraries\r\nDrive technical excellence in performance optimization and system reliability across the Neuron ecosystem\r\nMentor team members and provide technical leadership across multiple work streams\r\nDrive architectural decisions that impact the entire Neuron serving stack\r\nCollaborate with customers, product owners, and engineering teams to define technical strategy\r\nAuthor technical documentation, design proposals, and architectural guidelines\r\nA day in the life:\r\nLeading critical technical initiatives while mentoring team members\r\nCollaborating with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities\r\nLeading design reviews and architectural discussions\r\nDebugging complex performance issues across the stack in collaboration with the compiler and runtime teams\r\nMentoring junior engineers on system design and model optimization across model enablement teams\r\nDriving technical decisions that shape the future of Neuron's inference stack\r\nAbout the team:\r\nThe inference model enablement team releases its models in the vLLM Neuron plugin.","company":"Amazon","rawCompany":"amazon","city":"Cupertino","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-08T01:24:12.522Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron","description":"Sr. Software Development EngineerWe develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.\r\nAs a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.\r\nKey job responsibilities:\r\nDeliver high-performance models using distributed inference libraries\r\nDrive technical excellence in performance optimization and system reliability across the Neuron ecosystem\r\nMentor team members and provide technical leadership across multiple work streams\r\nDrive architectural decisions that impact the entire Neuron serving stack\r\nCollaborate with customers, product owners, and engineering teams to define technical strategy\r\nAuthor technical documentation, design proposals, and architectural guidelines\r\nA day in the life:\r\nLeading critical technical initiatives while mentoring team members\r\nCollaborating with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities\r\nLeading design reviews and architectural discussions\r\nDebugging complex performance issues across the stack in collaboration with the compiler and runtime teams\r\nMentoring junior engineers on system design and model optimization across model enablement teams\r\nDriving technical decisions that shape the future of Neuron's inference stack\r\nAbout the team:\r\nThe inference model enablement team releases its models in the vLLM Neuron plugin.","datePosted":"2026-08-08T01:24:12.522Z","dateModified":"2026-08-08T01:24:12.522Z","hiringOrganization":{"@type":"Organization","name":"Amazon","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Cupertino","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"095902ffc9739bd6e7ed40ab"},"url":"https://jobsearcher.com/jobs/095902ffc9739bd6e7ed40ab"}}