JOBSEARCHER

Platform / Inference Optimization Engineer

Platform Engineer - Inference Optimization We build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization of SOTA serving frameworks and building a scalable, low-latency, cost-efficient AI platform across inference, RL training, and Kubernetes infrastructure. Responsibilities Deeply customize and optimize SOTA LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, and NVIDIA Dynamo at the source code level (PyTorch / C++ / CUDA). Redesign and optimize scheduler logic (priority scheduling, preemption, request reordering, prefill/decode interleaving). Develop and tune custom CUDA kernels. Perform cross-layer debugging from Python serving stack down to C++/CUDA kernels to eliminate performance bottlenecks. Build and maintain GPU-native Kubernetes infrastructure, including custom controllers/operators, topology-aware scheduling, and multi-region deployments. Optimize high-performance networking (RDMA / InfiniBand / GPUDirect) for tensor-parallel and multi-node inference workloads. Build distributed training and RL infrastructure (GRPO, RLHF), including rollout services, reward model serving, and hybrid train/infer scheduling. Establish inference observability systems (TTFT, TPOT, KV cache hit rate, GPU fragmentation, scheduler queue depth) and use metrics to drive automated scaling and SLA protection. Requirements Strong proficiency in Python and C++, with the ability to read, modify, and contribute to large open-source projects (e.g., PyTorch, vLLM, TensorRT-LLM). Hands‑on experience customizing or optimizing at least one SOTA LLM serving framework. Deep understanding of LLM inference optimizations, including quantization, speculative decoding, continuous batching, and KV cache management. Strong knowledge of GPU architecture and hardware‑aware performance tuning. Production experience with Kubernetes, Docker, and large‑scale distributed systems. Experience operating cloud‑based GPU infrastructure across regions and environments. Experience with monitoring systems (Prometheus, Grafana) and cost‑aware resource optimization. Nice to Have Experience optimizing 8B-70B models with the goal of maximizing throughput. Familiarity with heterogeneous GPU clusters and hardware‑level performance characteristics. Open‑source contributions to inference acceleration projects (vLLM, SGLang, FlashAttention, etc.). Experience building LLM inference platforms from scratch. CUDA kernel development experience and deep understanding of GPU execution models. High‑performance networking experience (RDMA, InfiniBand, GPUDirect). Multi‑cloud or multi‑region distributed infrastructure experience. Compensation $200,000 - $500,000 total compensation (base + equity), depending on experience and impact. #J-18808-Ljbffr