Site Reliability Engineer SRE
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
A company is looking for a Site Reliability Engineer (SRE) to enhance stability and performance across Azure and GCP-based platforms.
Key Responsibilities
Manage incidents and lead root cause analysis (RCA) during 24/7 on-call rotations
Build and maintain observability dashboards, and perform deep-dive troubleshooting of application issues
Define SLOs, SLIs, and implement AI-Ops tools for proactive reliability management
Required Qualifications
5+ years in Site Reliability, DevOps, or Cloud Operations roles
Experience in application-level troubleshooting and performance analysis
Expertise in Azure and GCP cloud operations
Knowledge of Terraform, Ansible, and CI/CD pipelines
Familiarity with observability and AI-Ops tools