Site Reliability Engineer- REMOTE - Onsite Training
Overview
In this role you will strengthen and scale our reliability-focused platforms, partnering with product and development teams to improve performance and availability. You will own and evolve CI/CD pipelines, containerized deployments, and monitoring to support rapid growth. You’ll help define health signals (SLIs) and build automation that drives quality gates for production. This position blends hands-on engineering with technical leadership in a Digital organization focused on reliable, scalable software.
Compensation / Benefitsmedical, dental, and vision insurance with employer contributionflexible spending or health savings accountslife and AD&D insuranceshort and long term disabilitypaid time offemployee assistance program (EAP) and 401k with company match
ResponsibilitiesImplement and support CI/CD tools and pipelines across the organization (GitLab preferred)Monitor production and non-production systems and troubleshoot with Datadog and similar toolsMaintain and optimize containerized applications using Docker and related technologiesCollaborate with product and development teams to scale applications and infrastructure for high availabilityDefine and align Service Level Indicators (SLIs) with development teamsDevelop systems that increase site reliability and contribute to SRE competencyEvolve CI/CD pipelines with monitoring insights and automation for quality gatesBuild long-term automation using scripting/programming (Groovy, Shell, Python, Terraform, Java/JavaScript)Provide technical leadership for UI, APIs, and microservicesParticipate in scheduled on-call rotation
Key requirements8+ years in DevOps/SRE environments with strong CI/CD, containerization, and Kubernetes-based deployments5+ years using APM tools like Datadog or Dynatrace4+ years scripting in Unix/Linux shell and supporting NoSQL databases (Couchbase)3+ years with static code analysis tools (Checkmarx, SonarQube)Experience with Docker/Kubernetes, GitLab, Jenkins, Bamboo, and monitoring toolingBachelor in Computer Science or equivalent work experienceExperience with ServiceNow and Jira for change management (preferred)OWASP knowledge (preferred)strong problem-solving and diagnostic abilitiesleadership and mentorship in technical contextscontinuous improvement mindset and operational excellenceCI/CD pipelines (GitLab, Jenkins, Bamboo)Kubernetes and EKS deploymentsDocker and container orchestration