{"schemaVersion":"jobsearcher.job.v1","id":"02e037e0aaf41caa4f1daf38","url":"https://jobsearcher.com/jobs/02e037e0aaf41caa4f1daf38","canonicalUrl":"https://jobsearcher.com/jobs/02e037e0aaf41caa4f1daf38","title":"Senior ML Systems Engineer, GPU Performance","description":"Summary:High-growth AI infrastructure investment (with offices in SF and NYC) is looking to expand its ML performance team. The company supports mission-critical inference workloads for many of the world’s fastest-growing AI businesses and is building the infrastructure developers use to deploy open-source models in production.This role is focused on making large language model inference faster and more efficient. You will implement and productionize advanced optimization techniques across a wide range of model architectures, while working deep within the frameworks and runtimes that power modern AI inference.Our target profile will have strong backend or ML systems experience and a demonstrated ability to improve software performance. You should be comfortable debugging across Python, C++, GPU kernels, ML frameworks, and distributed production infrastructure.Qualifications:Inference Optimization: Experience with techniques such as quantization, speculative decoding, continuous batching, KV cache reuse, chunked prefill, or LoRA.ML Frameworks: Strong familiarity with PyTorch, TensorRT, TensorRT-LLM, vLLM, SGLang, or similar inference frameworks and runtimes.GPU Performance: Deep understanding of GPU architecture and experience diagnosing performance bottlenecks across compute, memory, and communication.Programming: Strong experience with Python, C++, or another systems-oriented programming language.Production ML Systems: Experience developing and deploying scalable AI/ML inference systems in production environments.Performance Engineering: Ability to investigate complex codebases, profile system behavior, and implement measurable improvements across different model architectures.Infrastructure: Experience with CUDA, Docker, Kubernetes, or distributed GPU environments is a strong plus.About Us:Greylock is an early-stage investor in hundreds of remarkable companies including Airbnb, LinkedIn, Dropbox, Workday, Cloudera, Facebook, Instagram, Roblox, Coinbase, Palo Alto Networks, among others. More can be found about us here: https://greylock.com/How We Work:We are full-time, salaried employees of Greylock and provide free candidate referrals/introductions to our active investments. We will contact anyone who looks like a potential match, requesting to schedule a call with you immediately.Due to the selective nature of this service and the volume of applicants we typically receive from our job postings, a follow-up email will not be sent until a match is identified with one of our investments.Please note: We are not recruiting for any roles within Greylock at this time. This job posting is for direct employment with a startup in our portfolio.","company":"Greylock Partners","rawCompany":"greylock partners","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-30T11:04:09.259Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior ML Systems Engineer, GPU Performance","description":"Summary:High-growth AI infrastructure investment (with offices in SF and NYC) is looking to expand its ML performance team. The company supports mission-critical inference workloads for many of the world’s fastest-growing AI businesses and is building the infrastructure developers use to deploy open-source models in production.This role is focused on making large language model inference faster and more efficient. You will implement and productionize advanced optimization techniques across a wide range of model architectures, while working deep within the frameworks and runtimes that power modern AI inference.Our target profile will have strong backend or ML systems experience and a demonstrated ability to improve software performance. You should be comfortable debugging across Python, C++, GPU kernels, ML frameworks, and distributed production infrastructure.Qualifications:Inference Optimization: Experience with techniques such as quantization, speculative decoding, continuous batching, KV cache reuse, chunked prefill, or LoRA.ML Frameworks: Strong familiarity with PyTorch, TensorRT, TensorRT-LLM, vLLM, SGLang, or similar inference frameworks and runtimes.GPU Performance: Deep understanding of GPU architecture and experience diagnosing performance bottlenecks across compute, memory, and communication.Programming: Strong experience with Python, C++, or another systems-oriented programming language.Production ML Systems: Experience developing and deploying scalable AI/ML inference systems in production environments.Performance Engineering: Ability to investigate complex codebases, profile system behavior, and implement measurable improvements across different model architectures.Infrastructure: Experience with CUDA, Docker, Kubernetes, or distributed GPU environments is a strong plus.About Us:Greylock is an early-stage investor in hundreds of remarkable companies including Airbnb, LinkedIn, Dropbox, Workday, Cloudera, Facebook, Instagram, Roblox, Coinbase, Palo Alto Networks, among others. More can be found about us here: https://greylock.com/How We Work:We are full-time, salaried employees of Greylock and provide free candidate referrals/introductions to our active investments. We will contact anyone who looks like a potential match, requesting to schedule a call with you immediately.Due to the selective nature of this service and the volume of applicants we typically receive from our job postings, a follow-up email will not be sent until a match is identified with one of our investments.Please note: We are not recruiting for any roles within Greylock at this time. This job posting is for direct employment with a startup in our portfolio.","datePosted":"2026-07-30T11:04:09.259Z","dateModified":"2026-07-30T11:04:09.259Z","hiringOrganization":{"@type":"Organization","name":"Greylock Partners","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"02e037e0aaf41caa4f1daf38"},"url":"https://jobsearcher.com/jobs/02e037e0aaf41caa4f1daf38"}}