{"schemaVersion":"jobsearcher.job.v1","id":"4b00cf755b62437bb4b64fb2","url":"https://jobsearcher.com/jobs/4b00cf755b62437bb4b64fb2","canonicalUrl":"https://jobsearcher.com/jobs/4b00cf755b62437bb4b64fb2","title":"Platform / Inference Optimization Engineer","description":"Platform Engineer - Inference Optimization\nWe build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization of SOTA serving frameworks and building a scalable, low-latency, cost-efficient AI platform across inference, RL training, and Kubernetes infrastructure.\n\nResponsibilities\n\nDeeply customize and optimize SOTA LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, and NVIDIA Dynamo at the source code level (PyTorch / C++ / CUDA).\n\nRedesign and optimize scheduler logic (priority scheduling, preemption, request reordering, prefill/decode interleaving).\n\nDevelop and tune custom CUDA kernels.\n\nPerform cross-layer debugging from Python serving stack down to C++/CUDA kernels to eliminate performance bottlenecks.\n\nBuild and maintain GPU-native Kubernetes infrastructure, including custom controllers/operators, topology-aware scheduling, and multi-region deployments.\n\nOptimize high-performance networking (RDMA / InfiniBand / GPUDirect) for tensor-parallel and multi-node inference workloads.\n\nBuild distributed training and RL infrastructure (GRPO, RLHF), including rollout services, reward model serving, and hybrid train/infer scheduling.\n\nEstablish inference observability systems (TTFT, TPOT, KV cache hit rate, GPU fragmentation, scheduler queue depth) and use metrics to drive automated scaling and SLA protection.\n\nRequirements\n\nStrong proficiency in Python and C++, with the ability to read, modify, and contribute to large open-source projects (e.g., PyTorch, vLLM, TensorRT-LLM).\n\nHands‑on experience customizing or optimizing at least one SOTA LLM serving framework.\n\nDeep understanding of LLM inference optimizations, including quantization, speculative decoding, continuous batching, and KV cache management.\n\nStrong knowledge of GPU architecture and hardware‑aware performance tuning.\n\nProduction experience with Kubernetes, Docker, and large‑scale distributed systems.\n\nExperience operating cloud‑based GPU infrastructure across regions and environments.\n\nExperience with monitoring systems (Prometheus, Grafana) and cost‑aware resource optimization.\n\nNice to Have\n\nExperience optimizing 8B-70B models with the goal of maximizing throughput.\n\nFamiliarity with heterogeneous GPU clusters and hardware‑level performance characteristics.\n\nOpen‑source contributions to inference acceleration projects (vLLM, SGLang, FlashAttention, etc.).\n\nExperience building LLM inference platforms from scratch.\n\nCUDA kernel development experience and deep understanding of GPU execution models.\n\nHigh‑performance networking experience (RDMA, InfiniBand, GPUDirect).\n\nMulti‑cloud or multi‑region distributed infrastructure experience.\n\nCompensation\n\n$200,000 - $500,000 total compensation (base + equity), depending on experience and impact.\n\n#J-18808-Ljbffr","company":"Kaon Prev Flowgpt","rawCompany":"kaon prev flowgpt","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:53:43.962Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Platform / Inference Optimization Engineer","description":"Platform Engineer - Inference Optimization\nWe build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization of SOTA serving frameworks and building a scalable, low-latency, cost-efficient AI platform across inference, RL training, and Kubernetes infrastructure.\n\nResponsibilities\n\nDeeply customize and optimize SOTA LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, and NVIDIA Dynamo at the source code level (PyTorch / C++ / CUDA).\n\nRedesign and optimize scheduler logic (priority scheduling, preemption, request reordering, prefill/decode interleaving).\n\nDevelop and tune custom CUDA kernels.\n\nPerform cross-layer debugging from Python serving stack down to C++/CUDA kernels to eliminate performance bottlenecks.\n\nBuild and maintain GPU-native Kubernetes infrastructure, including custom controllers/operators, topology-aware scheduling, and multi-region deployments.\n\nOptimize high-performance networking (RDMA / InfiniBand / GPUDirect) for tensor-parallel and multi-node inference workloads.\n\nBuild distributed training and RL infrastructure (GRPO, RLHF), including rollout services, reward model serving, and hybrid train/infer scheduling.\n\nEstablish inference observability systems (TTFT, TPOT, KV cache hit rate, GPU fragmentation, scheduler queue depth) and use metrics to drive automated scaling and SLA protection.\n\nRequirements\n\nStrong proficiency in Python and C++, with the ability to read, modify, and contribute to large open-source projects (e.g., PyTorch, vLLM, TensorRT-LLM).\n\nHands‑on experience customizing or optimizing at least one SOTA LLM serving framework.\n\nDeep understanding of LLM inference optimizations, including quantization, speculative decoding, continuous batching, and KV cache management.\n\nStrong knowledge of GPU architecture and hardware‑aware performance tuning.\n\nProduction experience with Kubernetes, Docker, and large‑scale distributed systems.\n\nExperience operating cloud‑based GPU infrastructure across regions and environments.\n\nExperience with monitoring systems (Prometheus, Grafana) and cost‑aware resource optimization.\n\nNice to Have\n\nExperience optimizing 8B-70B models with the goal of maximizing throughput.\n\nFamiliarity with heterogeneous GPU clusters and hardware‑level performance characteristics.\n\nOpen‑source contributions to inference acceleration projects (vLLM, SGLang, FlashAttention, etc.).\n\nExperience building LLM inference platforms from scratch.\n\nCUDA kernel development experience and deep understanding of GPU execution models.\n\nHigh‑performance networking experience (RDMA, InfiniBand, GPUDirect).\n\nMulti‑cloud or multi‑region distributed infrastructure experience.\n\nCompensation\n\n$200,000 - $500,000 total compensation (base + equity), depending on experience and impact.\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:53:43.962Z","dateModified":"2026-07-16T03:53:43.962Z","hiringOrganization":{"@type":"Organization","name":"Kaon Prev Flowgpt","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"4b00cf755b62437bb4b64fb2"},"url":"https://jobsearcher.com/jobs/4b00cf755b62437bb4b64fb2"}}