{"schemaVersion":"jobsearcher.job.v1","id":"6d12a4292e75fe75594f3afe","url":"https://jobsearcher.com/jobs/6d12a4292e75fe75594f3afe","canonicalUrl":"https://jobsearcher.com/jobs/6d12a4292e75fe75594f3afe","title":"Reinforcement Learning Engineer","description":"Reinforcement Learning Engineer - Remote\n\nBright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.\nThis is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.\n\nJob Title: Reinforcement Learning Engineer\nLocation: 100% Remote (U.S.)\nPosition Type: Full-time, Direct W2\nSalary Range: $100,000–$150,000 Annually\nExperience Required: 6+ years\n\nSponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.\n\nKey Responsibilities\nDesign and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.\nDevelop, calibrate, and maintain simulation environments suitable for large-scale agent training.\nImplement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.\nEngineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.\nApply offline RL and imitation learning techniques where exploration is costly or unsafe.\nUse RLHF, DPO, and related techniques for fine-tuning large language models when relevant.\nBuild scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.\nOptimize training stability and sample efficiency through algorithmic and engineering improvements.\nDesign rigorous evaluation protocols, including out-of-distribution and adversarial test cases.\nImplement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.\nCollaborate with applied scientists and product teams to identify high-value RL use cases.\nMonitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.\nDocument methodology, design decisions, and operational characteristics for internal stakeholders.\nStay current with RL research and translate promising techniques into production-ready solutions.\n\nRequired Qualifications\nMaster’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.\nSix or more years of combined RL research and engineering experience.\nStrong proficiency in Python and modern deep learning frameworks.\nHands-on experience with at least one major RL library or in-house RL stack.\nSolid understanding of probability, optimization, and the theoretical foundations of RL.\nExperience designing and tuning reward functions in non-trivial environments.\nFamiliarity with simulation environments and large-scale experience collection.\nExperience training neural network policies on GPU clusters.\nStrong written and verbal communication skills.\nTrack record of shipping or publishing impactful RL work.\n\nPreferred Qualifications\nExperience with RLHF for large language models.\nFamiliarity with multi-agent RL or hierarchical RL.\nExposure to robotics, control systems, or autonomous driving.\nPublications in RL or related research venues.\nOpen-source contributions to RL libraries or environments.\n\nHow to Apply\nWould you like to know more about this opportunity? For immediate consideration, please send your resume to venkat.r@bvteck.com or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.\n\nBright Vision Technologies is an Equal Opportunity Employer.\n\nEqual Employment Opportunity (EEO) Statement\nBright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.\nBV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.\nh3WCwMOk3x","company":"Brightvisiontechnologies","rawCompany":"brightvisiontechnologies","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-11T16:08:45.300Z","occupations":[{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541990","title":"All Other Professional, Scientific, and Technical Services","slug":"all-other-professional-scientific-and-technical-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Reinforcement Learning Engineer","description":"Reinforcement Learning Engineer - Remote\n\nBright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.\nThis is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.\n\nJob Title: Reinforcement Learning Engineer\nLocation: 100% Remote (U.S.)\nPosition Type: Full-time, Direct W2\nSalary Range: $100,000–$150,000 Annually\nExperience Required: 6+ years\n\nSponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.\n\nKey Responsibilities\nDesign and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.\nDevelop, calibrate, and maintain simulation environments suitable for large-scale agent training.\nImplement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.\nEngineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.\nApply offline RL and imitation learning techniques where exploration is costly or unsafe.\nUse RLHF, DPO, and related techniques for fine-tuning large language models when relevant.\nBuild scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.\nOptimize training stability and sample efficiency through algorithmic and engineering improvements.\nDesign rigorous evaluation protocols, including out-of-distribution and adversarial test cases.\nImplement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.\nCollaborate with applied scientists and product teams to identify high-value RL use cases.\nMonitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.\nDocument methodology, design decisions, and operational characteristics for internal stakeholders.\nStay current with RL research and translate promising techniques into production-ready solutions.\n\nRequired Qualifications\nMaster’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.\nSix or more years of combined RL research and engineering experience.\nStrong proficiency in Python and modern deep learning frameworks.\nHands-on experience with at least one major RL library or in-house RL stack.\nSolid understanding of probability, optimization, and the theoretical foundations of RL.\nExperience designing and tuning reward functions in non-trivial environments.\nFamiliarity with simulation environments and large-scale experience collection.\nExperience training neural network policies on GPU clusters.\nStrong written and verbal communication skills.\nTrack record of shipping or publishing impactful RL work.\n\nPreferred Qualifications\nExperience with RLHF for large language models.\nFamiliarity with multi-agent RL or hierarchical RL.\nExposure to robotics, control systems, or autonomous driving.\nPublications in RL or related research venues.\nOpen-source contributions to RL libraries or environments.\n\nHow to Apply\nWould you like to know more about this opportunity? For immediate consideration, please send your resume to venkat.r@bvteck.com or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.\n\nBright Vision Technologies is an Equal Opportunity Employer.\n\nEqual Employment Opportunity (EEO) Statement\nBright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.\nBV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.\nh3WCwMOk3x","datePosted":"2026-08-11T16:08:45.300Z","dateModified":"2026-08-11T16:08:45.300Z","hiringOrganization":{"@type":"Organization","name":"Brightvisiontechnologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6d12a4292e75fe75594f3afe"},"url":"https://jobsearcher.com/jobs/6d12a4292e75fe75594f3afe"}}