AIOps / Observability Engineer
Job Summary
We are seeking an experienced *AIOps / Observability Engineer* to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred.
Key Responsibilities
* Design, implement, and maintain enterprise observability and monitoring solutions.
* Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.
* Build and integrate APIs to connect monitoring, alerting, and automation platforms.
* Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.
* Automate operational tasks using scripting languages and infrastructure-as-code tools.
* Monitor production environments to ensure high availability, performance, and reliability.
* Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.
* Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.
* Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.
* Implement best practices for observability, logging, tracing, and performance monitoring.
* Participate in on-call production support and continuous service improvement initiatives.
Required Skills
* 8+ years of experience in AIOps, Observability, SRE, or Production Support.
* Strong experience with observability and monitoring platforms such as *Dynatrace, Splunk, AppDynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack*.
* Hands-on experience with REST APIs and API integrations.
* Experience with cloud platforms including *AWS, Azure, or Google Cloud Platform (GCP)*.
* Strong scripting and automation experience using *Python, Shell, PowerShell, Ansible, Terraform, or similar technologies*.
* Experience with incident management, troubleshooting, and root cause analysis.
* Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.
* Experience supporting mission-critical production environments.
* Excellent analytical, communication, and problem-solving skills.
Preferred Qualifications
* Banking or Financial Services domain experience.
* Experience with CI/CD pipelines and DevOps practices.
* Knowledge of Kubernetes, Docker, and container observability.
* Familiarity with ITSM tools such as ServiceNow.
* Experience implementing AI/ML-driven monitoring and predictive analytics.
Nice to Have
* Kubernetes/OpenShift administration.
* Kafka or messaging platform monitoring.
* Infrastructure as Code (Terraform/CloudFormation).
* Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.
Pay: $45.00 - $52.00 per hour
Work Location: In person