{"schemaVersion":"jobsearcher.job.v1","id":"37edffebc1e261345b1b282f","url":"https://jobsearcher.com/jobs/37edffebc1e261345b1b282f","canonicalUrl":"https://jobsearcher.com/jobs/37edffebc1e261345b1b282f","title":"Machine Learning Engineer, LLM Inference Optimization","description":"Job DescriptionAbout Us\\nGMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably.\\n\\nGMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.\\n\\nThis role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open-source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what's published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.\\n\\nYou will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.\\n\\nKey Responsibilities\\n\\nDrive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack.\\nDevelop next-generation optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform.\\nAdvance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery).\\nDrive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency.\\nBuild scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve.\\nProductionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost.\\nEngage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies.\\nCollaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud.\\n\\n\\nRequired Skills\\n\\nStrong hands-on experience with LLM inference systems and performance optimization on modern GPUs.\\nSolid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs.\\nExperience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton.\\nDeep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling.\\nDemonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management.\\nStrong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins.\\nProficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging.\\nClear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally.\\n\\n\\nPreferred Qualifications\\n\\n2+ years of hands-on experience in LLM inference optimization, ML systems optimization, or PhD degree in related areas.\\nTrack record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production.\\nSpecific track depth in one or more of:\\nSpeculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production.\\nQuantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking.\\nPD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling.\\nKV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads.\\nFamiliarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus.\\nOpen-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects.\\nExperience publishing technical blogs, case studies, or papers on inference optimization.\\n","company":"Gmi Cloud","rawCompany":"gmi cloud","city":"Mundelein","state":"IL","isRemote":false,"isActive":false,"createdAt":"2026-08-22T12:19:25.318Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer, LLM Inference Optimization","description":"Job DescriptionAbout Us\\nGMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably.\\n\\nGMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.\\n\\nThis role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open-source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what's published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.\\n\\nYou will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.\\n\\nKey Responsibilities\\n\\nDrive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack.\\nDevelop next-generation optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform.\\nAdvance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery).\\nDrive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency.\\nBuild scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve.\\nProductionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost.\\nEngage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies.\\nCollaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud.\\n\\n\\nRequired Skills\\n\\nStrong hands-on experience with LLM inference systems and performance optimization on modern GPUs.\\nSolid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs.\\nExperience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton.\\nDeep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling.\\nDemonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management.\\nStrong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins.\\nProficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging.\\nClear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally.\\n\\n\\nPreferred Qualifications\\n\\n2+ years of hands-on experience in LLM inference optimization, ML systems optimization, or PhD degree in related areas.\\nTrack record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production.\\nSpecific track depth in one or more of:\\nSpeculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production.\\nQuantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking.\\nPD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling.\\nKV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads.\\nFamiliarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus.\\nOpen-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects.\\nExperience publishing technical blogs, case studies, or papers on inference optimization.\\n","datePosted":"2026-08-22T12:19:25.318Z","dateModified":"2026-08-22T12:19:25.318Z","hiringOrganization":{"@type":"Organization","name":"Gmi Cloud","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Mundelein","addressRegion":"IL","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"37edffebc1e261345b1b282f"},"url":"https://jobsearcher.com/jobs/37edffebc1e261345b1b282f"}}