Senior ML Systems Engineer, GPU Performance
Summary:High-growth AI infrastructure investment (with offices in SF and NYC) is looking to expand its ML performance team. The company supports mission-critical inference workloads for many of the world’s fastest-growing AI businesses and is building the infrastructure developers use to deploy open-source models in production.This role is focused on making large language model inference faster and more efficient. You will implement and productionize advanced optimization techniques across a wide range of model architectures, while working deep within the frameworks and runtimes that power modern AI inference.Our target profile will have strong backend or ML systems experience and a demonstrated ability to improve software performance. You should be comfortable debugging across Python, C++, GPU kernels, ML frameworks, and distributed production infrastructure.Qualifications:Inference Optimization: Experience with techniques such as quantization, speculative decoding, continuous batching, KV cache reuse, chunked prefill, or LoRA.ML Frameworks: Strong familiarity with PyTorch, TensorRT, TensorRT-LLM, vLLM, SGLang, or similar inference frameworks and runtimes.GPU Performance: Deep understanding of GPU architecture and experience diagnosing performance bottlenecks across compute, memory, and communication.Programming: Strong experience with Python, C++, or another systems-oriented programming language.Production ML Systems: Experience developing and deploying scalable AI/ML inference systems in production environments.Performance Engineering: Ability to investigate complex codebases, profile system behavior, and implement measurable improvements across different model architectures.Infrastructure: Experience with CUDA, Docker, Kubernetes, or distributed GPU environments is a strong plus.About Us:Greylock is an early-stage investor in hundreds of remarkable companies including Airbnb, LinkedIn, Dropbox, Workday, Cloudera, Facebook, Instagram, Roblox, Coinbase, Palo Alto Networks, among others. More can be found about us here: https://greylock.com/How We Work:We are full-time, salaried employees of Greylock and provide free candidate referrals/introductions to our active investments. We will contact anyone who looks like a potential match, requesting to schedule a call with you immediately.Due to the selective nature of this service and the volume of applicants we typically receive from our job postings, a follow-up email will not be sent until a match is identified with one of our investments.Please note: We are not recruiting for any roles within Greylock at this time. This job posting is for direct employment with a startup in our portfolio.