{"schemaVersion":"jobsearcher.job.v1","id":"e9bac8038d5e92dce467a59b","url":"https://jobsearcher.com/jobs/e9bac8038d5e92dce467a59b","canonicalUrl":"https://jobsearcher.com/jobs/e9bac8038d5e92dce467a59b","title":"Java SRE","description":"Job Details:Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.Work on integrating the existing AI tools and should know major AI frameworks and libraries.Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.Cross-team coordination, incident triage and resolution, leadership and stakeholder management.Perform root cause analysis, identify recurring failure patterns, track corrective/preventive actions, and drive permanent fixes.Review production changes, support deployments, perform pre/post-change validations, monitor critical releases, and support rollback/recovery when required.Drive EMIM bridges for application impacts, troubleshoot issues, coordinate dependent teams, provide technical updates, and support faster service restoration.Primary Skills:JavaSRELang-chain, Langraph, RAG, MCPExperience with working on LLM's and integrating with the existing applicationsPython - FastAPICache - RedisSecondary Skills:Observability - ELK (Elastic/Kibana), Prometheus, GrafanaSoftware and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, KafkaReliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services and event-driven architecture","company":"Atos","rawCompany":"atos","city":"Phoenix","state":"AZ","isRemote":false,"isActive":false,"createdAt":"2026-10-01T22:54:24.824Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Java SRE","description":"Job Details:Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.Work on integrating the existing AI tools and should know major AI frameworks and libraries.Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.Cross-team coordination, incident triage and resolution, leadership and stakeholder management.Perform root cause analysis, identify recurring failure patterns, track corrective/preventive actions, and drive permanent fixes.Review production changes, support deployments, perform pre/post-change validations, monitor critical releases, and support rollback/recovery when required.Drive EMIM bridges for application impacts, troubleshoot issues, coordinate dependent teams, provide technical updates, and support faster service restoration.Primary Skills:JavaSRELang-chain, Langraph, RAG, MCPExperience with working on LLM's and integrating with the existing applicationsPython - FastAPICache - RedisSecondary Skills:Observability - ELK (Elastic/Kibana), Prometheus, GrafanaSoftware and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, KafkaReliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services and event-driven architecture","datePosted":"2026-10-01T22:54:24.824Z","dateModified":"2026-10-01T22:54:24.824Z","hiringOrganization":{"@type":"Organization","name":"Atos","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Phoenix","addressRegion":"AZ","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e9bac8038d5e92dce467a59b"},"url":"https://jobsearcher.com/jobs/e9bac8038d5e92dce467a59b"}}