{"schemaVersion":"jobsearcher.job.v1","id":"70f820ade95e51c7414afda1","url":"https://jobsearcher.com/jobs/70f820ade95e51c7414afda1","canonicalUrl":"https://jobsearcher.com/jobs/70f820ade95e51c7414afda1","title":"Java SRE Engineer","description":"Must have skills includes Application Support, Observability tool knowledge with distributed background with tech stack of Java/Python. Having knowledge or work experience in AI is preferred.Supporting applications in production, including incident response, on-call rotations, and post-incident reviewsApplying observability engineering to our applications — defining SLOs/SLIs/error budgets, building dashboards, and implementing alerting strategies to proactively detect system degradation before customers are impactedInvestigating and resolving production issues, including performance tuning and capacity planningBuilding automation to reduce toil and improve developer productivityDriving independent initiatives to improve platform reliability, developer experience, and operational maturitySite Reliability Engineer ResponsibilitiesProactively identify reliability risks and independently drive initiatives to address them before they become incidentsParticipate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviewsDefine and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across servicesScope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversightMake sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teamsAutomate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as TerraformA good fit will haveAt least 5 years of experience working in a professional environment as a Site Reliability Engineer (or a Software Engineer with some SRE responsibilities)Strong hands-on experience with observability — you understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work togetherParticipated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughoutComfortable taking ownership of initiatives or projects independently, from scoping through to delivery, without needing constant directionContributed to the design, build, and operation of cloud-native applicationsExperience with automating repeatable tasks and processes Build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levelsTechnologies we use includePython, Java, and Go are our primary server languagesOur browser applications are based on Angular and ReactCode lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on JenkinsInfrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containersDatadog is our primary observability platform — experience with Datadog APM, dashboards, monitors, and RUM is a plusInfrastructure is managed as code using Terraform","company":"Atos","rawCompany":"atos","city":"Phoenix","state":"AZ","isRemote":false,"isActive":false,"createdAt":"2026-04-28T10:45:09.274Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Java SRE Engineer","description":"Must have skills includes Application Support, Observability tool knowledge with distributed background with tech stack of Java/Python. Having knowledge or work experience in AI is preferred.Supporting applications in production, including incident response, on-call rotations, and post-incident reviewsApplying observability engineering to our applications — defining SLOs/SLIs/error budgets, building dashboards, and implementing alerting strategies to proactively detect system degradation before customers are impactedInvestigating and resolving production issues, including performance tuning and capacity planningBuilding automation to reduce toil and improve developer productivityDriving independent initiatives to improve platform reliability, developer experience, and operational maturitySite Reliability Engineer ResponsibilitiesProactively identify reliability risks and independently drive initiatives to address them before they become incidentsParticipate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviewsDefine and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across servicesScope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversightMake sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teamsAutomate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as TerraformA good fit will haveAt least 5 years of experience working in a professional environment as a Site Reliability Engineer (or a Software Engineer with some SRE responsibilities)Strong hands-on experience with observability — you understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work togetherParticipated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughoutComfortable taking ownership of initiatives or projects independently, from scoping through to delivery, without needing constant directionContributed to the design, build, and operation of cloud-native applicationsExperience with automating repeatable tasks and processes Build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levelsTechnologies we use includePython, Java, and Go are our primary server languagesOur browser applications are based on Angular and ReactCode lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on JenkinsInfrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containersDatadog is our primary observability platform — experience with Datadog APM, dashboards, monitors, and RUM is a plusInfrastructure is managed as code using Terraform","datePosted":"2026-04-28T10:45:09.274Z","dateModified":"2026-04-28T10:45:09.274Z","hiringOrganization":{"@type":"Organization","name":"Atos","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Phoenix","addressRegion":"AZ","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"70f820ade95e51c7414afda1"},"url":"https://jobsearcher.com/jobs/70f820ade95e51c7414afda1"}}