Senior ML Software Engineer, Data Plane
Overview
In this role you own the inference data plane for a custom ML accelerator, building high-performance, end-to-end solutions from model definition to distributed execution. You will implement and optimize compute kernels, validate architectures, and drive integration with serving frameworks. Working in a frontier-scale environment, you shape performance across the stack and help bring models from validation to production. You’ll collaborate across teams to improve latency, throughput, and reliability on evolving hardware.
ResponsibilitiesDevelop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performanceImplement and validate end-to-end LLM architectures (decoder-only, mixture-of-experts) from PyTorch to distributed hardware executionIntegrate accelerator backends into open-source ML serving frameworks (vLLM, PyTorch) including scheduler, memory management, and model parallelismBuild and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targetsProfile and optimize inference workloads, instrument critical paths, and drive latency/throughput improvements from simulation to hardware bring-upOwn features end-to-end from design through implementation, testing, and integration into the software stackContribute to CI/CD that gates model and kernel changes on correctness and performance regressionsMentor engineers, lead design reviews, and raise engineering standards across the team
Key requirementsBachelor’s degree in computer science or equivalent7+ years of full software development lifecycle experienceKnowledge of ML and LLM fundamentals including transformer architectures and inference lifecyclesKnowledge of computer architecture, OS, and parallel computingStrong proficiency in C/C++Strong Linux systems knowledgeExperience developing compute kernels for GPUs, DSPs, or custom acceleratorsProven track record delivering complex software features end-to-endMentoring and leadershipStrong collaboration with cross-functional teamsProblem-solving and optimization mindsetML frameworks including PyTorch, JAX, vLLM, SGLang, Dynamo, TorchXLA, TensorRTExperience deploying LLMs on GPUs/Neural accelerators or CUDA kernel developmentFamiliarity with speculative decoding, KV cache optimization, and LLM serving optimizations