Machine Learning Performance Engineer, Annapurna Labs
Overview
In this role you will help build and optimize ML performance across the AWS Neuron stack for Inferentia and Trainium. You will work with cross-functional teams in Tel Aviv to push end-to-end system performance, kernel optimization, and SDK enhancements. You’ll tackle challenging problems across the stack and shape the direction of a new core team. This is a hands-on role at the intersection of ML and systems with meaningful impact on large-scale AI infrastructure.
Compensation / Benefitsflexible work hoursmentorship and career growthinclusive team culturework-life balancediverse experiences encouragedcareer advancement resources
ResponsibilitiesSolve cross-stack performance challenges and design innovative software to improve service performance, durability, cost, and securityResearch and implement solutions to deliver optimal experiences for customersDesign, implement, test, deploy, and maintain high-performance software across the ML stackDevelop high-impact solutions for a large customer base and contribute to end-to-end performance gainsParticipate in design discussions and code reviews; communicate with internal and external stakeholdersWork cross-functionally to inform technical decisions and drive business outcomesOperate in a startup-like environment focusing on the most important work
Key requirementsB.S. in Computer Science (or related field)Proficiency in Python or C++ (Python preferred)Experience with TensorFlow, PyTorch, and/or JAX3+ years of non-internship software development experience3+ years of experience in performance optimizations for deep-learning models (LLMs, vision, etc.)cross-functional collaborationstrong problem-solving mindseteffective communication with stakeholdersPythonC++TensorFlow