JOBSEARCHER

AI Benchmark Auditor (Python & Evaluation)

HumanitappBrooklyn, NYL4 MidSeptember 21st, 2026
HumanitApp is seeking a role focused on evaluating benchmark tasks used to train and assess frontier AI models, focusing on quality and reproducibility. You will assess tasks, patches, and harnesses while delivering rubric-based feedback to ensure rigorous evaluation standards. This remote-friendly opportunity offers compensation ranging from $70 to $90 per hour, with workload and terms clarified during the official application process. #J-18808-Ljbffr