{"schemaVersion":"jobsearcher.job.v1","id":"c13d26f4b149f21d690f2f28","url":"https://jobsearcher.com/jobs/c13d26f4b149f21d690f2f28","canonicalUrl":"https://jobsearcher.com/jobs/c13d26f4b149f21d690f2f28","title":"Overwatch — Observability & Evaluation Engineer","description":"Job SummaryWe’re looking for an experienced Observability & Evaluation Engineer to help build the telemetry, evaluation, monitoring, and operational foundations required to safely release and operate AI agents in production.In this role, you’ll work at the intersection of LLM evaluation, AI agent observability, test automation, and production engineering, ensuring priority agent releases are measurable, reliable, and production-ready.Role: Observability & Evaluation EngineerFocus: LLMs | AI Agents | Evaluation | Observability | Python | Production OperationsWhat You’ll DoImplement telemetry, distributed tracing, and observability for LLM and AI agent workflows.Design and maintain LLM and agent evaluation suites to measure quality, accuracy, reliability, and performance.Build metrics, dashboards, alerts, and monitoring for production AI systems.Define and monitor Service Level Objectives (SLOs) and operational health indicators.Develop automated testing and evaluation frameworks using Python.Analyze prompt performance, model behavior, latency, cost, and quality metrics.Establish readiness criteria and provide production readiness evidence for priority agent releases.Create and maintain runbooks for incident response, troubleshooting, and operational support.Partner with AI/ML engineers, software engineers, product teams, and platform teams to improve agent reliability.Continuously identify opportunities to improve evaluation coverage, observability, and production resilience.Key Skills & ExperienceLLM / Generative AI evaluationAI Agent evaluation and testingObservability, telemetry & distributed tracingMetrics, dashboards & alertingSLOs / SLIs / production monitoringTest automation & quality engineeringPrompt and model performance analysisStrong Python programming skillsProduction operations & troubleshootingExperience working with cloud-native / distributed systemsWhat We’re Looking ForSomeone who can go beyond simply monitoring systems and answer critical questions such as:“Is our AI agent actually working well in production?”“How do we measure agent quality?”“Can we confidently release this model or prompt change?”“What happens when the agent starts degrading?”If you enjoy building the observability and evaluation layer behind production-grade AI agents, we'd love to hear from you!","company":"Codernation Technologies","rawCompany":"codernation technologies","city":"Charlotte","state":"NC","isRemote":false,"isActive":false,"createdAt":"2026-09-06T08:22:14.267Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Overwatch — Observability & Evaluation Engineer","description":"Job SummaryWe’re looking for an experienced Observability & Evaluation Engineer to help build the telemetry, evaluation, monitoring, and operational foundations required to safely release and operate AI agents in production.In this role, you’ll work at the intersection of LLM evaluation, AI agent observability, test automation, and production engineering, ensuring priority agent releases are measurable, reliable, and production-ready.Role: Observability & Evaluation EngineerFocus: LLMs | AI Agents | Evaluation | Observability | Python | Production OperationsWhat You’ll DoImplement telemetry, distributed tracing, and observability for LLM and AI agent workflows.Design and maintain LLM and agent evaluation suites to measure quality, accuracy, reliability, and performance.Build metrics, dashboards, alerts, and monitoring for production AI systems.Define and monitor Service Level Objectives (SLOs) and operational health indicators.Develop automated testing and evaluation frameworks using Python.Analyze prompt performance, model behavior, latency, cost, and quality metrics.Establish readiness criteria and provide production readiness evidence for priority agent releases.Create and maintain runbooks for incident response, troubleshooting, and operational support.Partner with AI/ML engineers, software engineers, product teams, and platform teams to improve agent reliability.Continuously identify opportunities to improve evaluation coverage, observability, and production resilience.Key Skills & ExperienceLLM / Generative AI evaluationAI Agent evaluation and testingObservability, telemetry & distributed tracingMetrics, dashboards & alertingSLOs / SLIs / production monitoringTest automation & quality engineeringPrompt and model performance analysisStrong Python programming skillsProduction operations & troubleshootingExperience working with cloud-native / distributed systemsWhat We’re Looking ForSomeone who can go beyond simply monitoring systems and answer critical questions such as:“Is our AI agent actually working well in production?”“How do we measure agent quality?”“Can we confidently release this model or prompt change?”“What happens when the agent starts degrading?”If you enjoy building the observability and evaluation layer behind production-grade AI agents, we'd love to hear from you!","datePosted":"2026-09-06T08:22:14.267Z","dateModified":"2026-09-06T08:22:14.267Z","hiringOrganization":{"@type":"Organization","name":"Codernation Technologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Charlotte","addressRegion":"NC","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c13d26f4b149f21d690f2f28"},"url":"https://jobsearcher.com/jobs/c13d26f4b149f21d690f2f28"}}