JOBSEARCHER

Site Reliability Engineer

Hello,We have multiple urgent openings for "Site Reliability Engineering". These are Hybrid rolesTitle: Site Reliability EngineeringLocation: Phoenix, Arizona (On-site from day 1, later hybrid 3 days per week)Only looking for candidates who can work onW2Strictly no C2C or third-party vendRequired QualificationsStrong Site Reliability Engineering experience supporting highly available, production-scale systems.Strong hands-on Google Cloud Platform (GCP) experience.Deep understanding and practical application of SLIs, SLOs, error budgets, operability, and reliability engineering principles.Strong experience with observability and instrumentation, including metrics, logging, tracing, alerting, and production diagnostics.Experience with production troubleshooting, incident response, root-cause analysis, and operational readiness.Experience with Terraform or another Infrastructure as Code technology.Experience developing and improving CI/CD pipelines and deployment practices.Proficiency with Python, Bash, or another scripting/automation language.Ability to identify systemic reliability issues and drive engineering solutions rather than primarily responding to operational incidents.Strong technical leadership, collaboration, and communication skills across application, platform, and engineering teams.Preferred QualificationsExperience supporting Kubernetes and GKE workloads in production.Experience with GitHub Actions.Experience with GCP Cloud Monitoring, OpenTelemetry, Prometheus, or Grafana.Experience with Istio or another service mesh.Experience with Apigee or another API gateway.Experience implementing canary, blue/green, or progressive delivery strategies.Experience improving engineering automation and automated testing practices.This is a hands-on technical leadership role with the opportunity to become a key technical authority for SRE practices across the organization.