Machine Learning Eval Engineer
Who is Recruiting from Scratch: Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire. Machine Learning Eval Engineer Location - San Francisco, CA (Onsite - 5 Days per Week) Compensation - $150,000 - $300,000 Base + Competitive Equity Visa - Visa Sponsorship Available (Case-by-Case) Company Stage - Series B (~$100M+ Raised) Industry - AI Infrastructure, Machine Learning, LLM Evaluation, Document Intelligence, Enterprise AI About the Company The company is building an AI-native document intelligence platform that enables enterprises to process, understand, and reason over complex unstructured documents at massive scale. Its platform combines proprietary document understanding models, frontier large language models, and enterprise-grade AI infrastructure to power highly accurate document workflows across industries. Already trusted by leading AI companies, Fortune-scale enterprises, and quantitative trading firms, the platform processes billions of documents while continuously improving model quality through sophisticated evaluation systems. Backed by world-class venture investors and built by an elite engineering team from companies including Stripe, Discord, Scale AI, and leading quantitative trading firms, the company is rapidly expanding its machine learning organization while maintaining an exceptionally high technical bar. As a Machine Learning Eval Engineer, you'll build the evaluation infrastructure that determines model quality, identifies failure modes, and drives improvements across production AI systems while working directly with machine learning, platform, and customer-facing teams. This is a rare opportunity to join one of the fastest-growing AI infrastructure companies where you'll directly influence how enterprise AI systems are measured, improved, and deployed at internet scale. What You'll DoDesign and build scalable evaluation systems for production LLM applicationsDevelop benchmarks, metrics, and automated evaluation pipelines measuring model qualityBuild workflows that identify failure modes across large-scale unstructured datasetsDesign statistical evaluation methodologies using precision, recall, and model quality metricsBuild internal tooling and lightweight applications for model visualization and evaluation analysisWork hands-on with enterprise documents including PDFs, spreadsheets, OCR outputs, and unstructured dataPartner closely with ML engineers to prioritize model improvements using evaluation insightsBuild customer-specific benchmarks demonstrating model performance across real-world workflowsDesign evaluation infrastructure supporting production AI systems operating at massive scaleCollaborate with GTM, Product, and Engineering teams to communicate model performancePrototype new evaluation techniques leveraging LLM-as-a-Judge methodologiesOwn evaluation systems from initial design through production deployment Ideal Candidate Background Experience Requirements1-5 years of experience in Machine Learning Engineering, Software Engineering, or ML InfrastructureStrong sweet spot around 2-4 years of experienceExperience building evaluation systems, ML tooling, or data infrastructure from zero-to-oneExperience working at high-bar technology companies, AI startups, quantitative firms, or leading research organizationsExperience working with production LLM applicationsExperience building customer-facing ML tooling or internal AI platformsStartup experience strongly preferredDemonstrated ownership of high-impact technical initiatives Technical RequirementsStrong Python engineering skillsDeep understanding of LLM evaluation methodologies including LLM-as-a-JudgeStrong prompt engineering experienceStrong understanding of precision, recall, statistical evaluation, and ML metricsExperience building evaluation pipelines or benchmarking systemsComfortable building lightweight web applications using Flask, TypeScript, or similar frameworksExperience working with unstructured data including documents, PDFs, OCR, or document extractionFamiliarity with AWS S3, OLAP systems, Tinybird, or analytics infrastructure preferredExperience working with Vision-Language Models or document AI preferredStrong debugging, experimentation, and software engineering fundamentals EducationBachelor's degree in Computer Science, Mathematics, Physics, Machine Learning, or related technical field preferredStrong academic background from a top engineering or quantitative program preferredFormal machine learning education or research experience preferred Soft SkillsStrong analytical thinkingHigh ownership mentalityExcellent communication skillsComfortable explaining technical concepts to non-technical stakeholdersComfortable operating in ambiguitySelf-directed and proactiveBias toward execution with technical precisionStartup mentalityStrong engineering craftsmanship Preferred BackgroundsAI infrastructure startupsLLM platform companiesDocument AI companiesMachine learning platform teamsQuantitative trading firmsAI research organizationsEarly-stage venture-backed startupsEvaluation infrastructure teamsData infrastructure organizationsEngineers building production AI systems Compensation & BenefitsBase Salary: $150,000 - $300,000Competitive Equity PackageDirect collaboration with ML and founding teamsSignificant ownership over evaluation infrastructureOpportunity to define model quality across enterprise AI systemsHigh-impact engineering roleExposure to cutting-edge LLM and document AI technologiesRapid career growth opportunitiesOnsite collaboration with a world-class engineering teamVisa Sponsorship Available (Case-by-Case) Why Join This is an opportunity to define how one of the industry's leading AI infrastructure platforms measures, improves, and scales model quality. You'll build evaluation systems, benchmarks, and tooling that directly influence production AI performance while collaborating closely with machine learning engineers, platform teams, and enterprise customers solving challenging real-world document intelligence problems. As one of the early ML Evaluation Engineers, you'll have outsized ownership over model quality infrastructure while helping build AI systems trusted by leading enterprises and frontier AI companies.