{"schemaVersion":"jobsearcher.job.v1","id":"47f9bef0d10cb88ec7daec2c","url":"https://jobsearcher.com/jobs/47f9bef0d10cb88ec7daec2c","canonicalUrl":"https://jobsearcher.com/jobs/47f9bef0d10cb88ec7daec2c","title":"MLOps Engineer","description":"About the team\n\nAI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.\n\nAbout the role\n\nThis role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.\n\nWhat you'll do\nDesign and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly\nBuild CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks\nOwn experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly\nMonitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale\nPartner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions\nEvaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform\nWhat you'll bring\nExperience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)\nPython or Go for pipeline orchestration and infrastructure tooling\nReal experience with experiment tracking and data versioning tools at production scale\nComfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales\nComfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment\nYou'll thrive here if\n\nYou treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.\n\nSPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.\n\n.css-16e3h2y{font-weight: normal;margin: 0;padding: 0;color: var(-color-on-surface, #000000);font-family: \"ripplingFontMedium\";font-size: 15px;letter-spacing: 0.25px;line-height: 22px;display: block;}\n\nThe pay range for this role is:\n\n.css-1xnp5ym{font-weight: normal;margin: 0;padding: 0;color: var(-color-on-surface, #000000);font-family: \"ripplingFontLight\";font-size: 15px;letter-spacing: 0.5px;line-height: 22px;display: block;}\n\n145,000 - 180,000 USD per year (Hybrid (San Francisco, California, US))","company":"Spreeai","rawCompany":"spreeai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-17T10:11:54.971Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"MLOps Engineer","description":"About the team\n\nAI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.\n\nAbout the role\n\nThis role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.\n\nWhat you'll do\nDesign and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly\nBuild CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks\nOwn experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly\nMonitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale\nPartner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions\nEvaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform\nWhat you'll bring\nExperience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)\nPython or Go for pipeline orchestration and infrastructure tooling\nReal experience with experiment tracking and data versioning tools at production scale\nComfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales\nComfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment\nYou'll thrive here if\n\nYou treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.\n\nSPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.\n\n.css-16e3h2y{font-weight: normal;margin: 0;padding: 0;color: var(-color-on-surface, #000000);font-family: \"ripplingFontMedium\";font-size: 15px;letter-spacing: 0.25px;line-height: 22px;display: block;}\n\nThe pay range for this role is:\n\n.css-1xnp5ym{font-weight: normal;margin: 0;padding: 0;color: var(-color-on-surface, #000000);font-family: \"ripplingFontLight\";font-size: 15px;letter-spacing: 0.5px;line-height: 22px;display: block;}\n\n145,000 - 180,000 USD per year (Hybrid (San Francisco, California, US))","datePosted":"2026-09-17T10:11:54.971Z","dateModified":"2026-09-17T10:11:54.971Z","hiringOrganization":{"@type":"Organization","name":"Spreeai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"47f9bef0d10cb88ec7daec2c"},"url":"https://jobsearcher.com/jobs/47f9bef0d10cb88ec7daec2c"}}