JOBSEARCHER

DevOps Engineer

About The RoleThe role is responsible for the availability, latency, performance, efficiency, and capacity management of a high-throughput cloud-native platform. This involves designing and implementing automated infrastructure pipelines that support rapid deployment cycles while maintaining strict uptime and reliability SLAs.The engineer will work closely with software engineering teams to containerize workloads, manage stateful and stateless services in Kubernetes, and build comprehensive observability frameworks that enable proactive system monitoring and rapid incident response.Key ResponsibilitiesDesign, provision, and maintain multi-region cloud infrastructure on AWS or GCP using Terraform and Infrastructure-as-Code best practicesManage and scale production-grade Kubernetes (EKS/GKE) clusters, including networking, ingress controllers, and IAM integrationsDevelop and optimize CI/CD pipelines using GitHub Actions, GitLab CI, or ArgoCD to automate build, test, and deployment processesImplement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and ELK/Datadog to monitor system health and latencyParticipate in a blameless on-call rotation, conducting post-mortems and implementing preventative automation to reduce operational toilCollaborate with security teams to enforce IAM roles, network security policies, and vulnerability scanning within build and runtime environmentsWhat We Are Looking For3–6 years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering supporting high-traffic production environmentsStrong proficiency with Infrastructure as Code (IaC) tools, specifically Terraform, and container orchestration with KubernetesSolid software engineering skills in at least one scripting or programming language, such as Python, Go, or BashDeep understanding of Linux systems administration, networking fundamentals (TCP/IP, DNS, VPCs), and cloud security best practicesExperience configuring and managing CI/CD tools and modern observability pipelines (Prometheus, Grafana, Datadog)Bonus: Experience with service meshes (Istio/Linkerd), GitOps workflows (ArgoCD/Flux), or managing distributed databases at scale