Site Reliability Engineer
About The RoleThe role owns the availability, latency, performance, efficiency, change management, and emergency response of critical production services.The team works alongside software engineers to design, build, and operate resilient distributed systems capable of scaling seamlessly under high load.Key ResponsibilitiesDesign, provision, and maintain cloud infrastructure using Terraform, Kubernetes, and AWS or GCPBuild robust CI/CD pipelines utilizing GitHub Actions or ArgoCD to ensure rapid, automated, and safe deploymentsEstablish comprehensive observability frameworks using Prometheus, Grafana, and Datadog for metric tracking, logging, and distributed tracingParticipate in an on-call rotation to troubleshoot and resolve production incidents, conducting rigorous post-mortem analysesAutomate manual operational tasks and infrastructure provisioning to eliminate toil and enforce reliability standardsCollaborate with development teams to conduct capacity planning, architectural reviews, and load testingWhat We Are Looking For4–7 years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering roles within high-scale environmentsDeep hands-on expertise with Kubernetes, containerization, and service mesh technologies in productionStrong proficiency in infrastructure-as-code tools such as Terraform, CloudFormation, or PulumiFluent in at least one scripting or programming language: Python, Go, or BashSolid understanding of networking fundamentals, TCP/IP, DNS, TLS, and load balancing architecturesBonus: Experience with chaos engineering methodologies and managing multi-region cloud deployments