{"schemaVersion":"jobsearcher.job.v1","id":"ec2836b7d58b0897cd9fb02d","url":"https://jobsearcher.com/jobs/ec2836b7d58b0897cd9fb02d","canonicalUrl":"https://jobsearcher.com/jobs/ec2836b7d58b0897cd9fb02d","title":"Reinforcement Learning Environment Engineer","description":"Reinforcement Learning Environment Engineer\r\nRL Environments; MLE; LLM Tasks; Difficulty Distribution; Remote Contractor; PST Overlap (=4h); Advanced English (C1/C2);\r\nWe're hiring RL Environments Engineers to design and build MLE/SWE environments that deliver high-quality, diverse tasks with minimal supervision. You will target a specific language model, meet a defined difficulty distribution, and deliver about one task every 10 hours. This is a remote contractor role with =4 hours overlap to PST and advanced English (C1/C2) required.\r\nAbout the company\r\nPreference Model is building the next generation of training data to power the future of AI. Today's models are powerful but fail to reach their potential across diverse use cases because so many of the tasks that we want to use these models for are outside of their training data distribution. Preference Model creates reinforcement learning environments that encapsulate real-world use cases, enabling AI systems to practice, adapt, and learn from feedback grounded in reality. We seek to bring the real world into distribution for the models.\r\nOur founding team has previous experience on Anthropic's data team building data infrastructure, tokenizers, and datasets behind the Claude model. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.\r\nThe company is backed by Tier 1 Silicon Valley VC.\r\nResponsibilities\r\nDesign and build MLE/SWE environments and diverse tasks.\r\nTarget a specified language model and satisfy the required difficulty distribution.\r\nDeliver ~1 task per 8-10 hours once onboarded.\r\nEdit tasks within 24 hours based on customer feedback.\r\nOnboard quickly and start delivering on day one with minimal supervision.\r\nRequirements\r\nWhat we're looking for (must-haves)\r\nStrong Python (engineering-quality, not notebook-only).\r\nHands-on LLM/GenAI work in production: you've shipped and operated real systems (not \"wrapped an API and called it AI\").\r\nStrong product/engineering ownership: comfortable building, fixing, and scaling end-to-end pipelines.\r\n=4 hours PST overlap and advanced English (C1/C2) for specs, reviews, and feedback.\r\nAbility to meet throughput expectations and respond quickly to feedback.\r\nStrong signals (nice-to-have, big plus)\r\nExperience in high-stakes or regulated domains (e.g., healthcare, finance, fraud/risk, safety-critical systems).\r\nExperience designing environments/tasks for RL and/or evaluations.\r\nExposure to RL / bandits / agentic systems (not required, but a strong signal).\r\nNot a fit if\r\nYou're primarily a prompt engineer without strong ML/engineering foundations.\r\nYou're a research-only / academic-only profile with little or no shipping/production ownership.\r\nYou've only built in notebooks or rely heavily on managed AutoML tools.\r\nWorking conditions\r\nhours/week - full time - need 4 hours overlap in the working hours with the team in Pacific time zone;\r\nDeliverables-driven; begin shipping on day one.\r\nConversion & relocation: Potential path to FTE and relocation to the Bay Area if performance and mutual fit align.\r\nContacts\r\nJ-18808-Ljbffr","company":"Open Data Science","rawCompany":"open data science","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-04-09T15:30:14.055Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Reinforcement Learning Environment Engineer","description":"Reinforcement Learning Environment Engineer\r\nRL Environments; MLE; LLM Tasks; Difficulty Distribution; Remote Contractor; PST Overlap (=4h); Advanced English (C1/C2);\r\nWe're hiring RL Environments Engineers to design and build MLE/SWE environments that deliver high-quality, diverse tasks with minimal supervision. You will target a specific language model, meet a defined difficulty distribution, and deliver about one task every 10 hours. This is a remote contractor role with =4 hours overlap to PST and advanced English (C1/C2) required.\r\nAbout the company\r\nPreference Model is building the next generation of training data to power the future of AI. Today's models are powerful but fail to reach their potential across diverse use cases because so many of the tasks that we want to use these models for are outside of their training data distribution. Preference Model creates reinforcement learning environments that encapsulate real-world use cases, enabling AI systems to practice, adapt, and learn from feedback grounded in reality. We seek to bring the real world into distribution for the models.\r\nOur founding team has previous experience on Anthropic's data team building data infrastructure, tokenizers, and datasets behind the Claude model. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.\r\nThe company is backed by Tier 1 Silicon Valley VC.\r\nResponsibilities\r\nDesign and build MLE/SWE environments and diverse tasks.\r\nTarget a specified language model and satisfy the required difficulty distribution.\r\nDeliver ~1 task per 8-10 hours once onboarded.\r\nEdit tasks within 24 hours based on customer feedback.\r\nOnboard quickly and start delivering on day one with minimal supervision.\r\nRequirements\r\nWhat we're looking for (must-haves)\r\nStrong Python (engineering-quality, not notebook-only).\r\nHands-on LLM/GenAI work in production: you've shipped and operated real systems (not \"wrapped an API and called it AI\").\r\nStrong product/engineering ownership: comfortable building, fixing, and scaling end-to-end pipelines.\r\n=4 hours PST overlap and advanced English (C1/C2) for specs, reviews, and feedback.\r\nAbility to meet throughput expectations and respond quickly to feedback.\r\nStrong signals (nice-to-have, big plus)\r\nExperience in high-stakes or regulated domains (e.g., healthcare, finance, fraud/risk, safety-critical systems).\r\nExperience designing environments/tasks for RL and/or evaluations.\r\nExposure to RL / bandits / agentic systems (not required, but a strong signal).\r\nNot a fit if\r\nYou're primarily a prompt engineer without strong ML/engineering foundations.\r\nYou're a research-only / academic-only profile with little or no shipping/production ownership.\r\nYou've only built in notebooks or rely heavily on managed AutoML tools.\r\nWorking conditions\r\nhours/week - full time - need 4 hours overlap in the working hours with the team in Pacific time zone;\r\nDeliverables-driven; begin shipping on day one.\r\nConversion & relocation: Potential path to FTE and relocation to the Bay Area if performance and mutual fit align.\r\nContacts\r\nJ-18808-Ljbffr","datePosted":"2026-04-09T15:30:14.055Z","dateModified":"2026-04-09T15:30:14.055Z","hiringOrganization":{"@type":"Organization","name":"Open Data Science","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"ec2836b7d58b0897cd9fb02d"},"url":"https://jobsearcher.com/jobs/ec2836b7d58b0897cd9fb02d"}}