Senior Software Engineer — AI Evaluation & Benchmarks
About The RoleWhat if the code you write could determine how smart the next generation of AI truly is? We're looking for experienced Software Engineers to design and build the coding benchmarks and data pipelines used to evaluate frontier AI models — the systems that decide whether an AI can actually reason, debug, and write production-quality software.This is high-impact, technically demanding work at the intersection of software engineering and AI research. You'll work with large codebases, multiple programming languages, and scalable infrastructure to create evaluation systems that push the boundaries of what AI can do.This is a fully remote contract role. If you thrive in fast-paced engineering environments and want your work to directly shape the trajectory of AI — this is the role.Organization: AlignerrType: Hourly ContractLocation: RemoteContract Length: 3 MonthsCommitment: Full-time availability preferredWhat You'll DoDesign and implement coding benchmarks used to evaluate frontier AI models across real-world programming tasksBuild and maintain scalable data pipelines for AI evaluation workflowsAnalyze AI-generated code for correctness, reliability, and edge-case failuresCreate structured evaluation scenarios that rigorously test reasoning, debugging, and code qualityWork with large code repositories and multi-language environmentsCollaborate on systems that improve how AI models understand and generate softwareProvide detailed technical feedback on model performance and failure patternsContribute to the design of evaluation frameworks that set industry standardsWho You Are4+ years of professional software engineering experience — this is non-negotiableExperience working at a high-growth tech company or top-tier software organizationExpert proficiency in Python — you write clean, performant, well-tested Python codeHands-on experience with code repositories and working in large, complex codebasesProven experience designing and implementing LLM coding benchmarks and data pipelinesTrack record of working in high-performance engineering environments with large-scale products or platformsStrong command of version control systems (Git) and modern development workflowsBilingual or native English speaker with strong written communication skillsSelf-directed, technically rigorous, and comfortable operating with autonomyWhat Makes a Perfect MatchCandidates with these additional qualifications have the highest chance of success:Senior or Lead-level engineering profiles with a history of technical ownershipBachelor's or Master's degree in Computer Science, Machine Learning, or a related field — or equivalent professional experienceProficiency in one or more additional languages: JavaScript, Go, C++, or other relevant languagesExperience with CI/CD pipelines and writing robust unit tests (pytest, Mocha, JUnit)Background in security engineering or significant open-source contributionsFamiliarity with AI/ML evaluation methodologies or model benchmarkingWhy Join UsWork on cutting-edge AI evaluation projects alongside world-class research teamsFully remote — work from anywhere with a reliable internet connectionYour benchmarks directly influence how the most advanced AI systems in the world are measured and improvedFreelance autonomy with meaningful, high-stakes engineering workCollaborate with a global community of elite engineers and researchersPotential for contract extension and ongoing engagement as new evaluation challenges emerge