ML Software Engineer, Data Plane
Overview
In this role you will own the design and implementation of the inference data plane for large-scale AI models, optimizing run-time performance on custom hardware. You will work across model execution, memory management, data movement, and serving integration to deliver production-grade inference for frontier-scale models. You’ll collaborate with cross-functional teams to validate architectures end-to-end and drive performance from simulation through hardware bring-up. This is a ground-up effort in a fast-evolving hardware and software landscape with a focus on impactful, scalable ML inference.
Compensation / Benefitssign-on paymentsRSUshealth insurance (medical, dental, vision)401(k) matchingpaid time offparental leave
ResponsibilitiesDevelop and optimize compute kernels for a custom ML accelerator to deliver production-level LLM inference performanceImplement and validate LLM architectures end-to-end from PyTorch model definition to distributed execution on custom hardwareIntegrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch) including scheduler extensions, memory management, and model parallelismBuild and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targetsProfile and optimize inference workloads, identify bottlenecks, instrument critical paths, and push latency/throughput improvements from simulation to hardware bring-upOwn features end-to-end from design to implementation, testing, and integration into the broader software stackContribute to CI/CD pipelines ensuring correctness and performance of model and kernel changes
Key requirementsBachelor's degree or equivalent4+ years of full software development life cycle experienceKnowledge of computer architecture, operating systems, and parallel computingKnowledge of Linux fundamentalsStrong proficiency in C/C++Experience developing compute kernels for GPUs, DSPs, or custom acceleratorsProven track record of owning and delivering complex software features end-to-endownership and accountabilityproblem-solvingeffective communicationC/C++compute kernel development for GPUs/DSPs/custom acceleratorsLinux