SRE Architect
*Job Summary*
We are seeking an experienced *SRE Architect* with 15+ years of experience to lead the design and implementation of highly available, scalable, and resilient enterprise platforms. The ideal candidate will have deep expertise in Site Reliability Engineering, cloud platforms, Kubernetes, Infrastructure as Code, CI/CD automation, and observability. The candidate should also possess strong Banking or Financial Services experience and be capable of driving enterprise reliability initiatives while collaborating with cross-functional teams.
*Key Responsibilities*
* Design and architect enterprise-scale, highly available, fault-tolerant cloud platforms.
* Lead Site Reliability Engineering (SRE) initiatives to improve platform reliability, scalability, and operational excellence.
* Design and implement cloud infrastructure on AWS, Azure, or GCP.
* Architect Kubernetes/OpenShift container platforms and containerized application deployments.
* Develop Infrastructure as Code (IaC) solutions using Terraform, CloudFormation, and Ansible.
* Design and optimize enterprise CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
* Build enterprise observability solutions using Splunk, Datadog, Dynatrace, Prometheus, Grafana, ELK, or New Relic.
* Lead Incident Management, Problem Management, Root Cause Analysis (RCA), and major production incident resolution.
* Design automation frameworks, self-healing capabilities, and operational resilience solutions.
* Collaborate with Development, Infrastructure, Security, and Business teams to drive platform reliability.
* Establish SRE best practices, governance, SLIs, SLOs, and error budgets across enterprise platforms.
* Mentor engineering teams and provide technical leadership on reliability engineering initiatives.
*Required Skills*
* 15+ years of experience in Site Reliability Engineering (SRE) and Production Engineering.
* Strong experience with AWS, Azure, or GCP cloud platforms.
* Hands-on expertise in Kubernetes, Docker, and OpenShift.
* Strong Infrastructure as Code experience using Terraform, CloudFormation, and Ansible.
* Experience building enterprise CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
* Strong expertise in Splunk, Datadog, Dynatrace, Prometheus, Grafana, ELK, or New Relic.
* Strong knowledge of Incident Management, Problem Management, RCA, and Service Availability.
* Strong scripting skills using Python, Bash, or Go.
* Experience implementing automation, self-healing systems, and operational resilience.
* Excellent communication, leadership, and stakeholder management skills.
*Mandatory Requirements*
* Banking or Financial Services domain experience.
* Experience designing enterprise-scale, highly available, fault-tolerant platforms.
* Proven experience leading architecture and reliability initiatives in enterprise environments.
Pay: $65.00 - $70.00 per hour
Work Location: In person