JOBSEARCHER

AI Benchmark Quality Engineer for Task Evaluation

ObsidianMillbrae, CAL4 MidSeptember 9th, 2026
Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level tasks, reference patches, test harnesses, and grading integrity, and provide rubric-based written feedback. Responsibilities include identifying gaps in tooling, ensuring task reproducibility across environments, and delivering concrete, rubric-based assessments to guide model training and #J-18808-Ljbffr