Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Overview
In this role you will help advance AWS Neuron by enabling and optimizing distributed inference on Inferentia and Trainium. You will work across frameworks, kernels, and compiler/runtime layers to maximize performance for large models and GenAI workloads. You’ll collaborate with customers and open-source ecosystems to scale AI acceleration and push the boundaries of ML systems. This is a hands-on, startup-like environment focused on impactful infrastructure challenges and continuous learning.
Compensation / BenefitsHealth insurance (medical, dental, vision)401(k) matchingpaid time offparenteral leaveRSUs and sign-on incentivescomprehensive benefits package
ResponsibilitiesDesign, develop, and optimize ML models and frameworks for deployment on custom ML hardware acceleratorsParticipate in the ML system development lifecycle including distributed architecture, performance profiling, and production deploymentBuild infrastructure to analyze and onboard diverse modelsDesign and implement high-performance kernels and features leveraging Neuron architectureAnalyze and optimize system-level performance across Neuron generationsConduct profiling to identify bottlenecks and apply optimizations (fusion, tiling, scheduling)Perform comprehensive testing including unit and end-to-end model validation with CI/CDCollaborate with customers to enable and optimize ML models on AWS acceleratorsWork with cross-functional teams (compiler, runtime, framework, hardware) on optimization techniquesEngage in design discussions, code reviews, and stakeholder communicationsDrive automation, metrics, and root-cause analysis for software defectsContribute to future architecture designs and open-source integration
Key requirementsBachelor's degree in computer science or equivalent3+ years of professional software development experience3+ years of design/architecture experience for systemsFundamentals of ML and LLMs, with optimization experienceProficiency in C++ and PythonStrong understanding of system performance, memory management, and parallel computingDebugging, profiling, and applying best software engineering practices in large-scale systemsCross-functional collaborationMentorship and knowledge sharingProblem-solving and debugging under pressurePyTorch familiarity (preferred)JIT compilation and AOT tracing familiarityCUDA kernels or low-level kernels familiarity