Neuron Runtime Software Development Engineer , Neuron Runtime
Overview
In this role you will build and maintain high-performance runtime libraries and drivers for AWS Neuron, enabling efficient ML workloads on Trainium and Inferentia. You will drive the full development life cycle of the Neuron Runtime and related components, collaborating with cross-functional teams to improve performance insights for customers. You will advance profiling capabilities and multi-framework support, helping customers optimize ML workloads across accelerators. This is a chance to shape scalable, reliable software at the core of AWS ML infrastructure.
Compensation / Benefitshealth insurance401(k) matchingRSUspaid time offparental leaveflexible working hours
ResponsibilitiesDevelop and maintain high-performance runtime libraries and drivers for ML applications and AI acceleratorsDesign, develop, and deploy Neuron Runtime and related componentsEnhance profiling capabilities to optimize AI workloads across hardware platformsImprove performance of ML kernels and ML frameworksManage full development life cycle of Neuron Runtime ensuring scalability and reliabilityCollaborate with cross-functional teams to ensure C++ compiler emits useful performance informationDrive innovations to support multiple frameworks (PyTorch, JAX, XLA)
Key requirements3+ years of non-internship software development experience2+ years of non-internship design or architecture experience for new or existing systemsExperience with distributed systems focusing on high availability and fault tolerancecross-functional collaborationcommunication and on-call readinessproblem-solving under pressurehigh-performance runtime librariesdrivers for ML acceleratorsdistributed systems design and reliability