{"schemaVersion":"jobsearcher.job.v1","id":"966182d69cdab05e330d4ac4","url":"https://jobsearcher.com/jobs/966182d69cdab05e330d4ac4","canonicalUrl":"https://jobsearcher.com/jobs/966182d69cdab05e330d4ac4","title":"Software Engineer - ML Infrastructure","description":"About Us\nWe're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.\n\nRole Overview\nWe’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.\n\nSitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.\n\nKey Responsibilities\nPartner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change\nBuild a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.\nBuild high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.\nPartner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.\nContribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.\nBuild the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.\nQualifications\n6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production\nHave 2+ years of experience building ML infrastructure or systems in production\nStrong Python skills and expertise in PyTorch or JAX\nExperience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research\nExperience building infrastructure or platforms specifically for research or machine learning workflows\nDeep experience building and operating Kubernetes and cloud infrastructure at scale\nExperience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient\nPrior experience as a technical lead or mentor for other engineers\nPreferred Qualifications\nExperience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy\nExperience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems\nExperience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production\nExperience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.\nExperience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.\nExperience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows\nFamiliarity with vision-language models (VLMs) or multimodal architectures\nThe anticipated annual base salary for this position is up to $250,000. This range does not include any other compensation components or other benefits for which an individual may be eligible. The actual base salary offered depends on a variety of factors, which may include as applicable, the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the location in which the applicant lives and/or from which they will be performing the job.","company":"Epsilon Labs","rawCompany":"epsilon labs","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-06T16:57:53.089Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer - ML Infrastructure","description":"About Us\nWe're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.\n\nRole Overview\nWe’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.\n\nSitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.\n\nKey Responsibilities\nPartner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change\nBuild a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.\nBuild high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.\nPartner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.\nContribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.\nBuild the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.\nQualifications\n6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production\nHave 2+ years of experience building ML infrastructure or systems in production\nStrong Python skills and expertise in PyTorch or JAX\nExperience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research\nExperience building infrastructure or platforms specifically for research or machine learning workflows\nDeep experience building and operating Kubernetes and cloud infrastructure at scale\nExperience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient\nPrior experience as a technical lead or mentor for other engineers\nPreferred Qualifications\nExperience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy\nExperience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems\nExperience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production\nExperience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.\nExperience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.\nExperience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows\nFamiliarity with vision-language models (VLMs) or multimodal architectures\nThe anticipated annual base salary for this position is up to $250,000. This range does not include any other compensation components or other benefits for which an individual may be eligible. The actual base salary offered depends on a variety of factors, which may include as applicable, the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the location in which the applicant lives and/or from which they will be performing the job.","datePosted":"2026-08-06T16:57:53.089Z","dateModified":"2026-08-06T16:57:53.089Z","hiringOrganization":{"@type":"Organization","name":"Epsilon Labs","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"966182d69cdab05e330d4ac4"},"url":"https://jobsearcher.com/jobs/966182d69cdab05e330d4ac4"}}