AI-Driven DevOps Engineer & Model Evaluator
Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and CI/CD tools. This sprint-based role involves evaluating bugs, edge cases, and failure modes with professional engineering judgment, applying real-world reliability concepts to production-scale systems.