{"schemaVersion":"jobsearcher.job.v1","id":"f77219a8553e6a8017d0ad57","url":"https://jobsearcher.com/jobs/f77219a8553e6a8017d0ad57","canonicalUrl":"https://jobsearcher.com/jobs/f77219a8553e6a8017d0ad57","title":"Remote Software Engineer - Python/Typescript","description":"About Turing\nTuring is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.\n\nAbout the Role\nWe’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.\n\nRather than primarily building production applications, you’ll work with coding agents across real-world repositories and assess the quality of their work. You’ll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.\n\nThink of the coding agent as another engineer whose work you’re reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?\n\nWhat You’ll Do\nEvaluate AI-generated code and solutions across real-world software repositories\nReview agent behavior, tool usage, and code changes for correctness and quality\nIdentify technical errors, weak approaches, and recurring model failure modes\nCompare model outputs and explain why one solution is better than another\nCreate and refine rubrics and evaluation criteria for coding tasks\nProduce high-quality evaluation and preference data used to improve coding models\nBuild and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows\nSynthesize findings from data work into clear write-ups, updates, and recommendations for the team\nCollaborate closely with researchers and engineers to translate qualitative judgment into scalable processes\nShare clear, actionable findings with AI researchers and engineers\n\nWhat We’re Looking For\n5+ years of hands-on software engineering experience\nStrong proficiency in Python, TypeScript/JavaScript, Go, or another major production language\nExperience working in substantial real-world codebases\nStrong code-review skills and technical judgment\nAbility to clearly explain why an implementation is correct, incorrect, or could be improved\nStrong written communication\nExperience using modern LLMs or AI coding tools\nExperience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.\n\nEngagement Details\nCompensation: Market rate; please provide a specific hourly rate expectation\nAvailability: 40 hours/week preferred, with at least 6 hours of Pacific Time overlap\nType: Independent contractor\nDuration: Approximately 3 months\nStart: As soon as possible\n\nEvaluation Process\nAI interview (~25 minutes)\nPractical code/AI evaluation exercise (~30 minutes)\nHiring manager interview (~20 minutes)\n\nThe practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.","company":"Turing","rawCompany":"turing","isRemote":true,"isActive":false,"createdAt":"2026-09-23T10:20:44.303Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1251.00","title":"Computer Programmers","slug":"computer-programmers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Remote Software Engineer - Python/Typescript","description":"About Turing\nTuring is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.\n\nAbout the Role\nWe’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.\n\nRather than primarily building production applications, you’ll work with coding agents across real-world repositories and assess the quality of their work. You’ll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.\n\nThink of the coding agent as another engineer whose work you’re reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?\n\nWhat You’ll Do\nEvaluate AI-generated code and solutions across real-world software repositories\nReview agent behavior, tool usage, and code changes for correctness and quality\nIdentify technical errors, weak approaches, and recurring model failure modes\nCompare model outputs and explain why one solution is better than another\nCreate and refine rubrics and evaluation criteria for coding tasks\nProduce high-quality evaluation and preference data used to improve coding models\nBuild and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows\nSynthesize findings from data work into clear write-ups, updates, and recommendations for the team\nCollaborate closely with researchers and engineers to translate qualitative judgment into scalable processes\nShare clear, actionable findings with AI researchers and engineers\n\nWhat We’re Looking For\n5+ years of hands-on software engineering experience\nStrong proficiency in Python, TypeScript/JavaScript, Go, or another major production language\nExperience working in substantial real-world codebases\nStrong code-review skills and technical judgment\nAbility to clearly explain why an implementation is correct, incorrect, or could be improved\nStrong written communication\nExperience using modern LLMs or AI coding tools\nExperience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.\n\nEngagement Details\nCompensation: Market rate; please provide a specific hourly rate expectation\nAvailability: 40 hours/week preferred, with at least 6 hours of Pacific Time overlap\nType: Independent contractor\nDuration: Approximately 3 months\nStart: As soon as possible\n\nEvaluation Process\nAI interview (~25 minutes)\nPractical code/AI evaluation exercise (~30 minutes)\nHiring manager interview (~20 minutes)\n\nThe practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.","datePosted":"2026-09-23T10:20:44.303Z","dateModified":"2026-09-23T10:20:44.303Z","hiringOrganization":{"@type":"Organization","name":"Turing","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f77219a8553e6a8017d0ad57"},"url":"https://jobsearcher.com/jobs/f77219a8553e6a8017d0ad57"}}