{"schemaVersion":"jobsearcher.job.v1","id":"a454699923fd7fe5ebebb69d","url":"https://jobsearcher.com/jobs/a454699923fd7fe5ebebb69d","canonicalUrl":"https://jobsearcher.com/jobs/a454699923fd7fe5ebebb69d","title":"Senior Software Engineer Model Hosting- remote","description":"Senior Software Engineer - Model Hosting\n\nThe Model Hosting team sits at the heart of our infrastructure, working deeply within our bare metal and Kubernetes stack to ensure every bit of compute is used to its fullest. If you're excited about squeezing maximum performance out of hardware and building the systems that serve large language models at scale, we'd love to talk.\n\nAbout the Role\nYou'll help design and operate the platform that serves our models in production. Our stack spans several languages, each chosen for the job it does best: Go powers our gateway layer, Rust handles fast real-time decision-making, and Python supports model dependencies that require it. Kubernetes orchestrates all model deployments across the company.\n\nYour work will touch some of the most interesting problems in high-performance model serving, including:\n• RDMA technologies (InfiniBand and RoCE) across multiple nodes\n• Disaggregated serving and KV cache offloading\n• Quantization, speculative-decoding, MTP, and other inference performance optimizations\n• Working with modern inference engines such as SGLang, vLLM, and others\n\nWhat We're Looking For\n\nWe are primarily looking for someone with 1-3 years of experience hosting models (a fairly new field), but who also has 5-10+ years of broader software engineering experience.\n\n• Experience serving LLMs in production\n• Strong programming skills in Python and/or Go\n• A track record of designing and operating highly scalable, highly available distributed services\n• Working knowledge of InfiniBand, RoCE, or other high-performance networking\n\nNice to Have\n• Contributions to vLLM, SGLang, TRT-LLM, or other NVIDIA ecosystem open-source projects\n• Deep understanding of Kubernetes\n• Experience developing in Rust\n• Familiarity with storage over RDMA","company":"Calance","rawCompany":"calance","isRemote":true,"isActive":false,"createdAt":"2026-09-10T11:20:21.506Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Software Engineer Model Hosting- remote","description":"Senior Software Engineer - Model Hosting\n\nThe Model Hosting team sits at the heart of our infrastructure, working deeply within our bare metal and Kubernetes stack to ensure every bit of compute is used to its fullest. If you're excited about squeezing maximum performance out of hardware and building the systems that serve large language models at scale, we'd love to talk.\n\nAbout the Role\nYou'll help design and operate the platform that serves our models in production. Our stack spans several languages, each chosen for the job it does best: Go powers our gateway layer, Rust handles fast real-time decision-making, and Python supports model dependencies that require it. Kubernetes orchestrates all model deployments across the company.\n\nYour work will touch some of the most interesting problems in high-performance model serving, including:\n• RDMA technologies (InfiniBand and RoCE) across multiple nodes\n• Disaggregated serving and KV cache offloading\n• Quantization, speculative-decoding, MTP, and other inference performance optimizations\n• Working with modern inference engines such as SGLang, vLLM, and others\n\nWhat We're Looking For\n\nWe are primarily looking for someone with 1-3 years of experience hosting models (a fairly new field), but who also has 5-10+ years of broader software engineering experience.\n\n• Experience serving LLMs in production\n• Strong programming skills in Python and/or Go\n• A track record of designing and operating highly scalable, highly available distributed services\n• Working knowledge of InfiniBand, RoCE, or other high-performance networking\n\nNice to Have\n• Contributions to vLLM, SGLang, TRT-LLM, or other NVIDIA ecosystem open-source projects\n• Deep understanding of Kubernetes\n• Experience developing in Rust\n• Familiarity with storage over RDMA","datePosted":"2026-09-10T11:20:21.506Z","dateModified":"2026-09-10T11:20:21.506Z","hiringOrganization":{"@type":"Organization","name":"Calance","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a454699923fd7fe5ebebb69d"},"url":"https://jobsearcher.com/jobs/a454699923fd7fe5ebebb69d"}}