JOBSEARCHER

Software Engineer (AI Technical Advisor)

We are helping our client find exceptional Software Engineer - Technical Advisors who want to spend a focused period doing something most engineers never get to do: find out exactly where frontier models break on real engineering work.You will take on hard problems across production-grade codebases, then work out precisely where and why model-generated solutions fall short. You will also build the hard problems those models are tested on. Your judgment about what separates good engineering from plausible engineering is the entire point of the role.What you'll doEvaluate Agent Sessions: Analyze how AI coding agents behave over entire coding sessions—tracking what the agent investigates, verifies, assumes, and leaves undone.Review Model-Written PRs: Audit model-generated pull requests against real production repositories, documenting every identified issue alongside its severity and detailed technical rationale.Build Evaluation Benchmarks: Design and construct hard, container-based problems that serve as rigorous test benchmarks for frontier models.Collaborate with AI Researchers: Work directly alongside AI researchers on frontier problems, producing clear written analyses that explain root-cause failures and boundary-condition gaps.Maintain High Written Rationale Standards: Author original, clear technical rationale for all evaluations (all written deliverables must be independently authored without AI text generation, though AI tools are welcomed for codebase exploration and running test suites).Ideal candidate profileAre comfortable dropping into unfamiliar codebases and languages you don't use every day - the repos change weekly.Are comfortable with containers and reproducing results locally.Can explain your reasoning as clearly as you can write the solution.Are more interested in why a solution is right than in shipping it quickly.Want to spend a focused period on this rather than committing to a permanent role.Daily tasksSolve difficult engineering problems across real, production-grade codebases.Identify where model-generated code fails, and articulate precisely why.Work directly with Anthropic researchers on problems they are actively investigating.Hold a technical bar that others build on.Required skillsHave deep, demonstrated expertise as a software engineer - we care about the depth of your judgment, not your years.Write code that other strong engineers learn from.Have reviewed a lot of other people's code, and are known for catching what CI and the author both missed.