{"schemaVersion":"jobsearcher.job.v1","id":"cd1110eae939b41c74f80273","url":"https://jobsearcher.com/jobs/cd1110eae939b41c74f80273","canonicalUrl":"https://jobsearcher.com/jobs/cd1110eae939b41c74f80273","title":"Eval360 - Error Analysis Engineer","description":"About the Institute of Foundation Models\nThe Institute of Foundation Models is a dedicated research lab focused on building, understanding, using, and risk‑managing foundation models. Our mission is to advance AI research, support the next generation of AI builders, and develop impactful systems that improve how frontier models are trained, evaluated, deployed, and governed.\n\nAs part of our team, you will work closely with researchers, machine learning engineers, data scientists, software engineers, and product teams on some of the most important challenges in AI development. You will contribute to systems that help measure model quality, identify failure modes, and improve the reliability, safety, and readiness of model releases.\n\nEval360 - Error Analysis Engineer\nWe are looking for an Eval360 - Error Analysis Engineer to help build, improve, and operate Eval360, an evaluation service that serves as a quality gate for AI models. This person will focus specifically on error analysis: understanding where models fail, why they fail, how those failures should be categorized, and how evaluation systems can better detect, measure, and prevent these issues before models are released.\n\nYou will collaborate with researchers, machine learning engineers, product managers, data scientists, and platform teams to develop AI evaluation applications and internal tools based on next‑generation AI research. You will be part of a cross‑functional team responsible for the full software development lifecycle, from requirements gathering and system design to implementation, deployment, monitoring, debugging, documentation, and continuous improvement.\n\nThe ideal candidate is comfortable working across the stack, including front‑end interfaces for reviewing errors, back‑end evaluation pipelines, data analysis workflows, model evaluation infrastructure, databases, dashboards, and APIs. This person should have strong software engineering skills, excellent analytical judgment, and the ability to turn ambiguous model failures into structured insights that improve evaluation quality.\n\nCollaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.\n\nBuild and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.\n\nPerform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.\n\nDevelop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.\n\nDesign and implement client‑side and server‑side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.\n\nDevelop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.\n\nBuild and maintain back‑end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.\n\nTest software to ensure responsiveness, correctness, reliability, and efficiency across evaluation workflows.\n\nTroubleshoot, debug, and upgrade evaluation systems, including identifying issues in data processing, evaluation metrics, model output handling, job orchestration, and user‑facing analysis tools.\n\nCreate and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.\n\nWrite clear technical documentation for Eval360 systems, error taxonomies, evaluation workflows, debugging procedures, and user‑facing tools.\n\nWork with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure‑mode visibility.\n\nKeep track of new development tools, evaluation frameworks, model analysis methods, data quality techniques, and architectures relevant to AI evaluation systems.\n\nContribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.\n\nHelp ensure Eval360 produces reliable, interpretable, and actionable signals for model quality gates.\n\nContribute to research publications, technical reports, internal knowledge sharing, and external presentations where appropriate.\n\nContribute to intellectual property and thought leadership in AI evaluation, error analysis, model quality measurement, and evaluation infrastructure.\n\nPerform all other duties as reasonably directed by the line manager that are aligned with these functional objectives.\n\nEducation\n\nBachelor's degree in Computer Science, Machine Learning, Data Science, Software Engineering, Statistics, or a related technical field required.\n\nMaster's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field preferred.\n\nExperience\n\nProven experience as a Software Engineer, Full Stack Developer, Machine Learning Evaluation Engineer, Data Scientist, AI Engineer, or similar role.\n\nExperience building software systems for AI, machine learning, data analysis, evaluation, annotation, experimentation, or model monitoring.\n\nExperience working with AI algorithms and the ability to develop systems that accommodate AI‑related requirements.\n\nExperience performing error analysis, model evaluation, data quality analysis, or failure‑mode investigation for machine learning or language model systems.\n\nExperience developing internal applications, dashboards, review tools, or web‑based workflows for technical users.\n\nFamiliarity with common software stacks, including front‑end frameworks, back‑end services, databases, APIs, and cloud or internal infrastructure.\n\nFamiliarity with GitHub, Git, CI/CD workflows, and collaborative software development practices.\n\nKnowledge of front‑end languages and libraries such as HTML, CSS, JavaScript, TypeScript, React, Angular, or similar technologies.\n\nKnowledge of back‑end languages and frameworks such as Python, Java, C#, Node.js, FastAPI, Flask, Django, or similar technologies.\n\nFamiliarity with databases such as MySQL, PostgreSQL, MongoDB, or other structured and unstructured data stores.\n\nFamiliarity with evaluation frameworks, experiment tracking systems, data pipelines, or machine learning infrastructure is strongly preferred.\n\nAbility to analyze complex model outputs and translate qualitative failures into structured, measurable categories.\n\nStrong problem‑solving and troubleshooting skills, especially for ambiguous technical issues involving models, data, metrics, and software systems.\n\nEffective communication and collaboration skills, with the ability to work across research, engineering, data, and product teams.\n\nStrong attention to detail and a high bar for evaluation quality, reliability, and interpretability.\n\nExperience with large language models, foundation models, multimodal models, or model evaluation systems.\n\nExperience designing or using error taxonomies, evaluation rubrics, benchmark datasets, human evaluation workflows, or automated grading systems.\n\nExperience with Python‑based data analysis tools such as pandas, NumPy, Jupyter, or similar.\n\nExperience with visualization or dashboarding tools for model quality analysis.\n\nExperience with distributed systems, job queues, workflow orchestration, or large‑scale data processing.\n\nExperience working in a research environment or with fast‑moving AI product and model teams.\n\nVisa Sponsorship\nThis position is eligible for visa sponsorship.\n\nBenefits Include\n\nComprehensive medical, dental, and vision benefits\n\nBonus\n\n401K plan\n\nGenerous paid time off, sick leave, and holidays\n\nPaid parental leave\n\nEmployee assistance program\n\nLife insurance and disability insurance\n\n#J-18808-Ljbffr","company":"Ifm Us","rawCompany":"ifm us","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T04:09:00.543Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"17-2112.02","title":"Validation Engineers","slug":"validation-engineers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Eval360 - Error Analysis Engineer","description":"About the Institute of Foundation Models\nThe Institute of Foundation Models is a dedicated research lab focused on building, understanding, using, and risk‑managing foundation models. Our mission is to advance AI research, support the next generation of AI builders, and develop impactful systems that improve how frontier models are trained, evaluated, deployed, and governed.\n\nAs part of our team, you will work closely with researchers, machine learning engineers, data scientists, software engineers, and product teams on some of the most important challenges in AI development. You will contribute to systems that help measure model quality, identify failure modes, and improve the reliability, safety, and readiness of model releases.\n\nEval360 - Error Analysis Engineer\nWe are looking for an Eval360 - Error Analysis Engineer to help build, improve, and operate Eval360, an evaluation service that serves as a quality gate for AI models. This person will focus specifically on error analysis: understanding where models fail, why they fail, how those failures should be categorized, and how evaluation systems can better detect, measure, and prevent these issues before models are released.\n\nYou will collaborate with researchers, machine learning engineers, product managers, data scientists, and platform teams to develop AI evaluation applications and internal tools based on next‑generation AI research. You will be part of a cross‑functional team responsible for the full software development lifecycle, from requirements gathering and system design to implementation, deployment, monitoring, debugging, documentation, and continuous improvement.\n\nThe ideal candidate is comfortable working across the stack, including front‑end interfaces for reviewing errors, back‑end evaluation pipelines, data analysis workflows, model evaluation infrastructure, databases, dashboards, and APIs. This person should have strong software engineering skills, excellent analytical judgment, and the ability to turn ambiguous model failures into structured insights that improve evaluation quality.\n\nCollaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.\n\nBuild and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.\n\nPerform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.\n\nDevelop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.\n\nDesign and implement client‑side and server‑side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.\n\nDevelop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.\n\nBuild and maintain back‑end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.\n\nTest software to ensure responsiveness, correctness, reliability, and efficiency across evaluation workflows.\n\nTroubleshoot, debug, and upgrade evaluation systems, including identifying issues in data processing, evaluation metrics, model output handling, job orchestration, and user‑facing analysis tools.\n\nCreate and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.\n\nWrite clear technical documentation for Eval360 systems, error taxonomies, evaluation workflows, debugging procedures, and user‑facing tools.\n\nWork with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure‑mode visibility.\n\nKeep track of new development tools, evaluation frameworks, model analysis methods, data quality techniques, and architectures relevant to AI evaluation systems.\n\nContribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.\n\nHelp ensure Eval360 produces reliable, interpretable, and actionable signals for model quality gates.\n\nContribute to research publications, technical reports, internal knowledge sharing, and external presentations where appropriate.\n\nContribute to intellectual property and thought leadership in AI evaluation, error analysis, model quality measurement, and evaluation infrastructure.\n\nPerform all other duties as reasonably directed by the line manager that are aligned with these functional objectives.\n\nEducation\n\nBachelor's degree in Computer Science, Machine Learning, Data Science, Software Engineering, Statistics, or a related technical field required.\n\nMaster's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field preferred.\n\nExperience\n\nProven experience as a Software Engineer, Full Stack Developer, Machine Learning Evaluation Engineer, Data Scientist, AI Engineer, or similar role.\n\nExperience building software systems for AI, machine learning, data analysis, evaluation, annotation, experimentation, or model monitoring.\n\nExperience working with AI algorithms and the ability to develop systems that accommodate AI‑related requirements.\n\nExperience performing error analysis, model evaluation, data quality analysis, or failure‑mode investigation for machine learning or language model systems.\n\nExperience developing internal applications, dashboards, review tools, or web‑based workflows for technical users.\n\nFamiliarity with common software stacks, including front‑end frameworks, back‑end services, databases, APIs, and cloud or internal infrastructure.\n\nFamiliarity with GitHub, Git, CI/CD workflows, and collaborative software development practices.\n\nKnowledge of front‑end languages and libraries such as HTML, CSS, JavaScript, TypeScript, React, Angular, or similar technologies.\n\nKnowledge of back‑end languages and frameworks such as Python, Java, C#, Node.js, FastAPI, Flask, Django, or similar technologies.\n\nFamiliarity with databases such as MySQL, PostgreSQL, MongoDB, or other structured and unstructured data stores.\n\nFamiliarity with evaluation frameworks, experiment tracking systems, data pipelines, or machine learning infrastructure is strongly preferred.\n\nAbility to analyze complex model outputs and translate qualitative failures into structured, measurable categories.\n\nStrong problem‑solving and troubleshooting skills, especially for ambiguous technical issues involving models, data, metrics, and software systems.\n\nEffective communication and collaboration skills, with the ability to work across research, engineering, data, and product teams.\n\nStrong attention to detail and a high bar for evaluation quality, reliability, and interpretability.\n\nExperience with large language models, foundation models, multimodal models, or model evaluation systems.\n\nExperience designing or using error taxonomies, evaluation rubrics, benchmark datasets, human evaluation workflows, or automated grading systems.\n\nExperience with Python‑based data analysis tools such as pandas, NumPy, Jupyter, or similar.\n\nExperience with visualization or dashboarding tools for model quality analysis.\n\nExperience with distributed systems, job queues, workflow orchestration, or large‑scale data processing.\n\nExperience working in a research environment or with fast‑moving AI product and model teams.\n\nVisa Sponsorship\nThis position is eligible for visa sponsorship.\n\nBenefits Include\n\nComprehensive medical, dental, and vision benefits\n\nBonus\n\n401K plan\n\nGenerous paid time off, sick leave, and holidays\n\nPaid parental leave\n\nEmployee assistance program\n\nLife insurance and disability insurance\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T04:09:00.543Z","dateModified":"2026-07-16T04:09:00.543Z","hiringOrganization":{"@type":"Organization","name":"Ifm Us","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"cd1110eae939b41c74f80273"},"url":"https://jobsearcher.com/jobs/cd1110eae939b41c74f80273"}}