{"schemaVersion":"jobsearcher.job.v1","id":"bcba2aca752cfce9871be3e8","url":"https://jobsearcher.com/jobs/bcba2aca752cfce9871be3e8","canonicalUrl":"https://jobsearcher.com/jobs/bcba2aca752cfce9871be3e8","title":"Research Engineer","description":"As a Research Engineer at Metis, you’ll work on building the next generation of autonomous post-training systems that leverage our Mantis platform. You’ll operate at the intersection of cutting-edge ML research and scalable engineering, designing, implementing, and deploying algorithms that improve how AI agents learn from feedback, synthetic data, and real-world interactions.\n\nYou’ll move seamlessly between papers and production, leading large-scale experiments, creating optimized training pipelines, and helping shape the future of post-training autonomy. You’ll have significant ownership, high compute budgets, and the mandate to push the state of the art in applied reinforcement and preference optimization.\n\nWhat You’ll Do\n\nResearch and help build an autonomous post-training agent leveraging the Mantis platform\n\nDesign and execute large-scale experiments on synthetic data generation and algorithmic architecture\n\nDevelop and refine methods for reinforcement learning, reward modeling, and human feedback integration\n\nCollaborate cross-functionally with Core and Platform Engineering to deploy and evaluate models in production settings\n\nPublish or contribute to leading-edge research in the post-training domain\n\nUse tooling and compute efficiently to iterate on experimental pipelines and accelerate research velocity\n\nRequirements\n\nDeep experience in machine learning, preferably reinforcement learning, post-training, or alignment research\n\nDemonstrated research contributions; ideally published papers (ICML, NeurIPS) or public implementations\n\nStrong proficiency in Python and ML frameworks (PyTorch, JAX, or TensorFlow)\n\nComfort with distributed training, high-throughput data pipelines, and large-scale experiment management\n\nAbility to reason independently, formulate hypotheses, and run experiments from idea to insight to product impact\n\nBase: $200,000-$1,000,000\n\nFull medical, dental, and vision\n\nWellness & L&D stipend\n\nBreakfast, lunch, and dinner provided (Unlimited Doordash)\n\n$25,000 housing stipend\n\n#J-18808-Ljbffr","company":"Metis","rawCompany":"metis","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:15:11.707Z","occupations":[{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"17-2199.08","title":"Robotics Engineers","slug":"robotics-engineers"}],"industries":[{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"},{"code":"541720","title":"Research and Development in the Social Sciences and Humanities","slug":"research-and-development-in-the-social-sciences-and-humanities"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Research Engineer","description":"As a Research Engineer at Metis, you’ll work on building the next generation of autonomous post-training systems that leverage our Mantis platform. You’ll operate at the intersection of cutting-edge ML research and scalable engineering, designing, implementing, and deploying algorithms that improve how AI agents learn from feedback, synthetic data, and real-world interactions.\n\nYou’ll move seamlessly between papers and production, leading large-scale experiments, creating optimized training pipelines, and helping shape the future of post-training autonomy. You’ll have significant ownership, high compute budgets, and the mandate to push the state of the art in applied reinforcement and preference optimization.\n\nWhat You’ll Do\n\nResearch and help build an autonomous post-training agent leveraging the Mantis platform\n\nDesign and execute large-scale experiments on synthetic data generation and algorithmic architecture\n\nDevelop and refine methods for reinforcement learning, reward modeling, and human feedback integration\n\nCollaborate cross-functionally with Core and Platform Engineering to deploy and evaluate models in production settings\n\nPublish or contribute to leading-edge research in the post-training domain\n\nUse tooling and compute efficiently to iterate on experimental pipelines and accelerate research velocity\n\nRequirements\n\nDeep experience in machine learning, preferably reinforcement learning, post-training, or alignment research\n\nDemonstrated research contributions; ideally published papers (ICML, NeurIPS) or public implementations\n\nStrong proficiency in Python and ML frameworks (PyTorch, JAX, or TensorFlow)\n\nComfort with distributed training, high-throughput data pipelines, and large-scale experiment management\n\nAbility to reason independently, formulate hypotheses, and run experiments from idea to insight to product impact\n\nBase: $200,000-$1,000,000\n\nFull medical, dental, and vision\n\nWellness & L&D stipend\n\nBreakfast, lunch, and dinner provided (Unlimited Doordash)\n\n$25,000 housing stipend\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:15:11.707Z","dateModified":"2026-07-16T03:15:11.707Z","hiringOrganization":{"@type":"Organization","name":"Metis","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"bcba2aca752cfce9871be3e8"},"url":"https://jobsearcher.com/jobs/bcba2aca752cfce9871be3e8"}}