Reliability Engineer
Overview
In this role, you will ensure the reliability, scalability, and health of EDAV’s Azure cloud environment, using Terraform to define infrastructure as code. You’ll collaborate with platform, development, and security teams to automate cloud operations, optimize AKS workloads, and improve incident response. The position emphasizes building scalable, observable systems and refining deployment practices. A meaningful hook is the chance to shape mission-critical cloud reliability at scale.
Compensation / BenefitsMedical, dental & vision401(k) retirement planLife Insurance (voluntary)Short and long-term disabilityHealth Spending Account (HSA)Time Off/Leave (PTO, Vacation or Sick Leave)
ResponsibilitiesDesign, implement, maintain, and troubleshoot production Azure infrastructure using TerraformSupport reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS)Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issuesAutomate cloud operations to reduce manual work and improve consistencyCollaborate with development and operations teams to enhance deployment and incident response practicesImplement and refine monitoring, alerting, dashboards, and operational reportingIdentify reliability risks and recommend improvements to cloud architecture and processesDocument infrastructure, procedures, troubleshooting guidance, and operational runbooks
Key requirementsdevopssremicrosoft azurekubernetesaksterraformcollaborativestrong problem-solvingclear communicationMicrosoft AzureKubernetesAzure Kubernetes Service (AKS)