{"schemaVersion":"jobsearcher.job.v1","id":"1a3c88e49302402f9fef35e9","url":"https://jobsearcher.com/jobs/1a3c88e49302402f9fef35e9","canonicalUrl":"https://jobsearcher.com/jobs/1a3c88e49302402f9fef35e9","title":"Founding Machine Learning - Eval Layer","description":"One Robot builds task-specific world models and an evaluation platform for robot manipulation policies.\nTraining end-to-end policies for robots is vibes-based today. Teams collect data, train, deploy on a real robot, find out what fails, collect more, retry. We replace the trial-and-error with rigorous validation that tells you where your policy will fail and what data to collect to fix it.\nRobotics can't industrialize without an evaluation layer. We're building it.\nWe're solving challenging technical problems around long-horizon autoregressive generation, world model controllability, and closing the sim-to-real gap. We work with real customer data, real failures, and real deployment pressure.\nWe're based in San Francisco, backed by Accel, YC, several exited founders, and engineering leaders at leading AI companies.\nWe're small and deliberately so. Everyone is an IC with deep ownership of a wide surface area. The culture is fast iteration and direct responsibility.\nHemanth Sarabu and Elton Shon co-founded One Robot after leading robot learning together at Industrial Next (YC W22), bringing experience from Google, NASA JPL, and Tesla.\nWe're building the evaluation layer to understand policy failure modes before they hit production. You'll own modeling work that makes the eval trustworthy.\nWhat you'll do:\nTrain evaluation models: Develop VLMs that classify and verify policy behavior.\nBuild confidence layers: Convert model outputs into trustworthy signals the customer can act on.\nImprove model grounding: Make the eval models reason accurately about physical and spatial scenes.\nBuild a self-improving eval layer: Develop data engine that makes the eval models sharper with each customer's deployments and corrections.\nRequirements:\nVery strong coding in Python and PyTorch.\nVLM/LLM training: Track record in training VLMs or LLMs.\nEvals experience: Developed and shipped evals for VLMs or LLMs.\nCompensation Range: $150K - $275K","company":"One Robot","rawCompany":"one robot","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-08T14:03:27.840Z","occupations":[{"code":"17-2199.08","title":"Robotics Engineers","slug":"robotics-engineers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"17-3024.01","title":"Robotics Technicians","slug":"robotics-technicians"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Machine Learning - Eval Layer","description":"One Robot builds task-specific world models and an evaluation platform for robot manipulation policies.\nTraining end-to-end policies for robots is vibes-based today. Teams collect data, train, deploy on a real robot, find out what fails, collect more, retry. We replace the trial-and-error with rigorous validation that tells you where your policy will fail and what data to collect to fix it.\nRobotics can't industrialize without an evaluation layer. We're building it.\nWe're solving challenging technical problems around long-horizon autoregressive generation, world model controllability, and closing the sim-to-real gap. We work with real customer data, real failures, and real deployment pressure.\nWe're based in San Francisco, backed by Accel, YC, several exited founders, and engineering leaders at leading AI companies.\nWe're small and deliberately so. Everyone is an IC with deep ownership of a wide surface area. The culture is fast iteration and direct responsibility.\nHemanth Sarabu and Elton Shon co-founded One Robot after leading robot learning together at Industrial Next (YC W22), bringing experience from Google, NASA JPL, and Tesla.\nWe're building the evaluation layer to understand policy failure modes before they hit production. You'll own modeling work that makes the eval trustworthy.\nWhat you'll do:\nTrain evaluation models: Develop VLMs that classify and verify policy behavior.\nBuild confidence layers: Convert model outputs into trustworthy signals the customer can act on.\nImprove model grounding: Make the eval models reason accurately about physical and spatial scenes.\nBuild a self-improving eval layer: Develop data engine that makes the eval models sharper with each customer's deployments and corrections.\nRequirements:\nVery strong coding in Python and PyTorch.\nVLM/LLM training: Track record in training VLMs or LLMs.\nEvals experience: Developed and shipped evals for VLMs or LLMs.\nCompensation Range: $150K - $275K","datePosted":"2026-08-08T14:03:27.840Z","dateModified":"2026-08-08T14:03:27.840Z","hiringOrganization":{"@type":"Organization","name":"One Robot","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"1a3c88e49302402f9fef35e9"},"url":"https://jobsearcher.com/jobs/1a3c88e49302402f9fef35e9"}}