Lead Site Reliability Engineer-9086
Overview
In this role you will provide technical leadership to design and build software and web applications, primarily using Python and Bash. You will manage container orchestration, cloud infrastructure, and CI/CD, shaping scalable, distributed systems for programmatic ad tech. You’ll support customer deployments, drive escalation responses, and mentor junior engineers. This remote-friendly position offers impact across production systems and open collaboration with cross-functional teams. You will help evolve tooling and processes to deliver reliable, high-performance advertising technology.
ResponsibilitiesArchitect and develop software and web apps using Python and BashManage Kubernetes and AWS ECS for container orchestration and cluster designAdminister AWS cloud infrastructure and managed servicesImplement metrics, logging, APM, and tracing with GrafanaHandle configuration management with Terraform, Ansible, and CloudFormationBuild CI/CD pipelines with Jenkins and AWS CodeDeployPerform Linux system administration and manage cloud-native networking (ALB, NLB, HAProxy, Nginx)Configure load balancing and messaging with Envoy Proxy, gRPC, and protobufAdminister Aerospike databasesUse ArgoCD and Argo Workflows for deployments and workflowsDevelop OpenRTB and VAST-based programmatic ad techAnalyze and troubleshoot large-scale distributed systemsPlan and lead scrum ceremonies to ensure timely deliverySupport production issues and provide root-cause fixesMentor junior engineers and act as technical escalation focal pointSupport customer deployments and expansion projectsLead data infrastructure monitoring and alerting improvementsGuide business and technical strategyInitialize and develop new technical toolsRemote work eligibility (100%)
Key requirementsBachelor’s degree and seven years of experience in software development and AWS infrastructureProficiency in Python and Bash scriptingExperience with AWS cloud infrastructure and managed servicesGrafana-based metrics visualization, logging, and APM tracingLinux systems administrationExperience with HAProxy and Nginx for load balancingFive years of configuration management with AnsibleCI/CD pipelines with JenkinsOne year of container orchestration with Kubernetes and AWS ECSExperience with Terraform and CloudFormationAWS CodeDeploy experienceNetworking with AWS ALB and NLBEnvoy Proxy, gRPC, and protobuf for messagingAerospike database administrationArgoCD and Argo Workflows for deploymentsOpenRTB and VAST protocolsMentoring and guiding junior engineersCross-functional collaborationProblem solving and root-cause analysisPythonBash scriptingKubernetes