Senior Software Engineer Model Hosting- remote
Senior Software Engineer - Model Hosting
The Model Hosting team sits at the heart of our infrastructure, working deeply within our bare metal and Kubernetes stack to ensure every bit of compute is used to its fullest. If you're excited about squeezing maximum performance out of hardware and building the systems that serve large language models at scale, we'd love to talk.
About the Role
You'll help design and operate the platform that serves our models in production. Our stack spans several languages, each chosen for the job it does best: Go powers our gateway layer, Rust handles fast real-time decision-making, and Python supports model dependencies that require it. Kubernetes orchestrates all model deployments across the company.
Your work will touch some of the most interesting problems in high-performance model serving, including:
• RDMA technologies (InfiniBand and RoCE) across multiple nodes
• Disaggregated serving and KV cache offloading
• Quantization, speculative-decoding, MTP, and other inference performance optimizations
• Working with modern inference engines such as SGLang, vLLM, and others
What We're Looking For
We are primarily looking for someone with 1-3 years of experience hosting models (a fairly new field), but who also has 5-10+ years of broader software engineering experience.
• Experience serving LLMs in production
• Strong programming skills in Python and/or Go
• A track record of designing and operating highly scalable, highly available distributed services
• Working knowledge of InfiniBand, RoCE, or other high-performance networking
Nice to Have
• Contributions to vLLM, SGLang, TRT-LLM, or other NVIDIA ecosystem open-source projects
• Deep understanding of Kubernetes
• Experience developing in Rust
• Familiarity with storage over RDMA