JOBSEARCHER

Machine Learning Engineer

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

Machine Learning Engineer- Applied ResearchLocation: HybridCompany: Pilotcrew AIType: Full-TimeExperience: 3-5 YearsAbout Pilotcrew AIPilotcrew AI builds infrastructure for AI Agent Evaluation. We benchmark large language models, run automated agent evaluations, power human-in-the-loop assessments, and host AI arenas for competitive testing.Our mission is to make AI agents measurable, reliable, and production-ready through structured, scalable evaluation systems.Role OverviewWe are hiring an Applied Research Engineer to bridge cutting‐edge AI research with production‐grade systems for evaluating LLMs and AI agents.In this role, you will read, interpret, and implement ideas from the latest research across large language models, multimodal systems, and agent architectures. You will translate these insights into scalable evaluation pipelines, new benchmarking methodologies, and improved model performance.You will work closely with engineering and product teams to turn research concepts into real‐world systems used for measuring, debugging, and improving AI agents.This is a research‐driven, execution‐heavy role requiring strong fundamentals, curiosity, and the ability to operate in a fast‐paced startup environment.Key ResponsibilitiesRead and synthesize research papers in LLMs, multimodal AI, and agent systemsImplement and adapt state‐of‐the‐art methods into production‐ready systemsDesign and improve evaluation methodologies (benchmarking, grading, scoring)Build experimental pipelines to test model behavior, robustness, and generalizationAnalyze model performance, failure modes, and edge casesDevelop novel metrics for reliability, reasoning quality, and tool usageContribute to adversarial testing and stress‐testing frameworksWork on multimodal systems (text, vision, tool interactions) where relevantCollaborate with engineering teams to productionize research ideasDocument findings and communicate insights clearly to technical stakeholdersRequired SkillsStrong Python programming skillsSolid foundation in machine learning and deep learningHands‐on experience with PyTorch or TensorFlowExperience working with LLMs, transformers, or multimodal modelsAbility to read and understand research papers and implement them effectivelyStrong analytical thinking and experimentation skillsExperience designing experiments and interpreting resultsFamiliarity with evaluation metrics and benchmarking methodologiesPreferred SkillsExperience with LLM evaluation, benchmarking, or alignmentFamiliarity with agent architectures (ReAct, tool‐calling, planning systems)Experience with multimodal models (vision‐language systems, CLIP, etc.)Knowledge of RLHF, reward modeling, or preference learningExperience with retrieval systems, search, or re‐rankingExposure to distributed systems or large‐scale experimentation pipelinesBackground in applied ML research (industry or academia)What We ValueStrong curiosity and research mindsetAbility to translate theory into practical systemsOwnership and bias toward executionComfort working with ambiguity and evolving problem spacesClear and structured technical communicationAbility to thrive in a fast‐paced startup environment with high ownershipWhy Join Pilotcrew AIWork on cutting‐edge problems in AI evaluation and reliabilityBridge research and real‐world AI systemsHigh ownership and autonomy in a fast‐moving teamOpportunity to shape how AI agents are evaluated at scale.Exposure to both research‐driven innovation and production systems#J-18808-Ljbffr