JOBSEARCHER

Senior AI/ML Engineer

UmatrMillbrae, CAL6 LeadAugust 28th, 2026
Job DescriptionSan Francisco | On-site | $150k–$275k + equityWe are working with a fast-growing AI startup building the operating brain for the supply chain. They’ve grown 10x in the last year with a small engineering team and are now building out the model layer underneath their production AI systems.They’re looking for their first dedicated ML Engineer to own models end-to-end, from raw data through to production. You’ll work with years of real-world operational data across 500k+ SKUs, building systems that directly impact how the business operates.This is not a research role, and it’s not an LLM-wrapper role. They’re looking for someone who can build, deploy and operate production ML systems - and take ownership when reality changes.What you'll ownBuild production forecasting models across messy, intermittent and seasonal demand, including cold-start SKUs, promotions, perishability and long-tail demandBuild datasets and fine-tune models using LoRA / PEFT, with rigorous evaluations determining what actually ships to productionBuild the representation layer that allows AI systems to reason across inconsistent products, vendors, pack sizes and units of measureOwn the infrastructure around those models, including deployment, versioning, monitoring, drift detection and automated retrainingBuild large-scale ML and data workloads using Python + SparkWork with AWS SageMaker, S3, Glue + Step FunctionsBuild production inference and evaluation infrastructureUse MLflow, Kubeflow or equivalent MLOps toolingContribute outside the model layer when needed, including enough TypeScript/React to work across the wider productThere are no handoffs. You’ll build the model, put it into production, monitor it and fix it when reality changes.What we're looking for5-7 years of experience building production ML systemsExperience building and maintaining time-series forecasting models serving production trafficHands-on experience with AWS SageMakerExperience fine-tuning LLMs using LoRA or PEFT on real datasetsExperience building systems backed by ontologies or knowledge graphsStrong experience engineering large-scale data pipelines with SparkExperience owning production models through deployment, monitoring, drift detection and retrainingStrong architecture skills, with the ability to explain and defend technical decisions in detailComfortable working across the full ML lifecycle rather than owning just one part of the processThey’re looking for someone who can talk about what happened after the model shipped - when it degraded, how you detected it, what it got wrong and what you changed.