{"schemaVersion":"jobsearcher.job.v1","id":"40aa5457d02982e959488ad0","url":"https://jobsearcher.com/jobs/40aa5457d02982e959488ad0","canonicalUrl":"https://jobsearcher.com/jobs/40aa5457d02982e959488ad0","title":"AIOps / Observability Engineer","description":"Job Summary\n\nWe are seeking an experienced *AIOps / Observability Engineer* to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred.\n\nKey Responsibilities\n\n* Design, implement, and maintain enterprise observability and monitoring solutions.\n* Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.\n* Build and integrate APIs to connect monitoring, alerting, and automation platforms.\n* Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.\n* Automate operational tasks using scripting languages and infrastructure-as-code tools.\n* Monitor production environments to ensure high availability, performance, and reliability.\n* Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.\n* Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.\n* Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.\n* Implement best practices for observability, logging, tracing, and performance monitoring.\n* Participate in on-call production support and continuous service improvement initiatives.\n\nRequired Skills\n\n* 8+ years of experience in AIOps, Observability, SRE, or Production Support.\n* Strong experience with observability and monitoring platforms such as *Dynatrace, Splunk, AppDynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack*.\n* Hands-on experience with REST APIs and API integrations.\n* Experience with cloud platforms including *AWS, Azure, or Google Cloud Platform (GCP)*.\n* Strong scripting and automation experience using *Python, Shell, PowerShell, Ansible, Terraform, or similar technologies*.\n* Experience with incident management, troubleshooting, and root cause analysis.\n* Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.\n* Experience supporting mission-critical production environments.\n* Excellent analytical, communication, and problem-solving skills.\n\nPreferred Qualifications\n\n* Banking or Financial Services domain experience.\n* Experience with CI/CD pipelines and DevOps practices.\n* Knowledge of Kubernetes, Docker, and container observability.\n* Familiarity with ITSM tools such as ServiceNow.\n* Experience implementing AI/ML-driven monitoring and predictive analytics.\n\nNice to Have\n\n* Kubernetes/OpenShift administration.\n* Kafka or messaging platform monitoring.\n* Infrastructure as Code (Terraform/CloudFormation).\n* Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.\n\nPay: $45.00 - $52.00 per hour\n\nWork Location: In person","company":"Cloudsecurityweb","rawCompany":"cloudsecurityweb","city":"Phoenix","state":"AZ","isRemote":false,"isActive":false,"createdAt":"2026-08-31T11:32:14.587Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AIOps / Observability Engineer","description":"Job Summary\n\nWe are seeking an experienced *AIOps / Observability Engineer* to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred.\n\nKey Responsibilities\n\n* Design, implement, and maintain enterprise observability and monitoring solutions.\n* Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.\n* Build and integrate APIs to connect monitoring, alerting, and automation platforms.\n* Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.\n* Automate operational tasks using scripting languages and infrastructure-as-code tools.\n* Monitor production environments to ensure high availability, performance, and reliability.\n* Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.\n* Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.\n* Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.\n* Implement best practices for observability, logging, tracing, and performance monitoring.\n* Participate in on-call production support and continuous service improvement initiatives.\n\nRequired Skills\n\n* 8+ years of experience in AIOps, Observability, SRE, or Production Support.\n* Strong experience with observability and monitoring platforms such as *Dynatrace, Splunk, AppDynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack*.\n* Hands-on experience with REST APIs and API integrations.\n* Experience with cloud platforms including *AWS, Azure, or Google Cloud Platform (GCP)*.\n* Strong scripting and automation experience using *Python, Shell, PowerShell, Ansible, Terraform, or similar technologies*.\n* Experience with incident management, troubleshooting, and root cause analysis.\n* Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.\n* Experience supporting mission-critical production environments.\n* Excellent analytical, communication, and problem-solving skills.\n\nPreferred Qualifications\n\n* Banking or Financial Services domain experience.\n* Experience with CI/CD pipelines and DevOps practices.\n* Knowledge of Kubernetes, Docker, and container observability.\n* Familiarity with ITSM tools such as ServiceNow.\n* Experience implementing AI/ML-driven monitoring and predictive analytics.\n\nNice to Have\n\n* Kubernetes/OpenShift administration.\n* Kafka or messaging platform monitoring.\n* Infrastructure as Code (Terraform/CloudFormation).\n* Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.\n\nPay: $45.00 - $52.00 per hour\n\nWork Location: In person","datePosted":"2026-08-31T11:32:14.587Z","dateModified":"2026-08-31T11:32:14.587Z","hiringOrganization":{"@type":"Organization","name":"Cloudsecurityweb","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Phoenix","addressRegion":"AZ","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"40aa5457d02982e959488ad0"},"url":"https://jobsearcher.com/jobs/40aa5457d02982e959488ad0"}}