{"schemaVersion":"jobsearcher.job.v1","id":"da57c28d4ebe94e2e1648453","url":"https://jobsearcher.com/jobs/da57c28d4ebe94e2e1648453","canonicalUrl":"https://jobsearcher.com/jobs/da57c28d4ebe94e2e1648453","title":"Software Development Engineer, SageMaker HyperPod Data Plane","description":"Overview\nIn this role you will lead the development of a next-generation AI compute platform optimized for LLMs and distributed training. You will work with ML scientists and customers to shape the roadmap and deliver scalable, robust ML services. You’ll own architecture decisions, drive best practices, and mentor junior engineers in a fast-paced, results-driven environment. This role sits at the intersection of platform engineering and customer-focused AI delivery, with a strong emphasis on Kubernetes-based orchestration, high-performance compute, and large-scale model training.\n\nCompensation / Benefitssign-on payments and RSUshealth insurance (medical, dental, vision)401(k) matchingpaid time offparental leaveflexible working hours\nResponsibilitiesDevelop innovative solutions for large-scale LLM training on a clustered node infrastructureBuild and maintain a performant and fully-managed service for training foundation modelsProfile and optimize distributed training to improve compute/network efficiencyServe as technical lead across project lifecycles from conception to deliveryHire and mentor junior engineersLeverage Kubernetes to build next-generation AI training platformCollaborate with ML scientists, internal teams, and open source communities (e.g., PyTorch, NVIDIA) to launch scalable products\nKey requirements1+ years of non-internship software development experienceProficiency in at least one programming languageExperience with multi-threaded asynchronous C++/Go developmentExperience with Kubernetes and resource orchestrationBackground in building scalable, high-performance systemsExperience in large language model training or HPC is a plusStrong communication skills and ability to work with customers and scientistscommunicationownershipanalytical leadershipC++GoKubernetes","company":"Amazonwebservices","rawCompany":"amazonwebservices","city":"San Jose","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-09-15T04:58:51.159Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1251.00","title":"Computer Programmers","slug":"computer-programmers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Development Engineer, SageMaker HyperPod Data Plane","description":"Overview\nIn this role you will lead the development of a next-generation AI compute platform optimized for LLMs and distributed training. You will work with ML scientists and customers to shape the roadmap and deliver scalable, robust ML services. You’ll own architecture decisions, drive best practices, and mentor junior engineers in a fast-paced, results-driven environment. This role sits at the intersection of platform engineering and customer-focused AI delivery, with a strong emphasis on Kubernetes-based orchestration, high-performance compute, and large-scale model training.\n\nCompensation / Benefitssign-on payments and RSUshealth insurance (medical, dental, vision)401(k) matchingpaid time offparental leaveflexible working hours\nResponsibilitiesDevelop innovative solutions for large-scale LLM training on a clustered node infrastructureBuild and maintain a performant and fully-managed service for training foundation modelsProfile and optimize distributed training to improve compute/network efficiencyServe as technical lead across project lifecycles from conception to deliveryHire and mentor junior engineersLeverage Kubernetes to build next-generation AI training platformCollaborate with ML scientists, internal teams, and open source communities (e.g., PyTorch, NVIDIA) to launch scalable products\nKey requirements1+ years of non-internship software development experienceProficiency in at least one programming languageExperience with multi-threaded asynchronous C++/Go developmentExperience with Kubernetes and resource orchestrationBackground in building scalable, high-performance systemsExperience in large language model training or HPC is a plusStrong communication skills and ability to work with customers and scientistscommunicationownershipanalytical leadershipC++GoKubernetes","datePosted":"2026-09-15T04:58:51.159Z","dateModified":"2026-09-15T04:58:51.159Z","hiringOrganization":{"@type":"Organization","name":"Amazonwebservices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"da57c28d4ebe94e2e1648453"},"url":"https://jobsearcher.com/jobs/da57c28d4ebe94e2e1648453"}}