{"schemaVersion":"jobsearcher.job.v1","id":"76bfcd7b8a26000a7070ab19","url":"https://jobsearcher.com/jobs/76bfcd7b8a26000a7070ab19","canonicalUrl":"https://jobsearcher.com/jobs/76bfcd7b8a26000a7070ab19","title":"Senior DevOps / Platform Reliability Engineer","description":"About Zingtree\nZingtree is the next-generation intelligent process automation platform reimagining customer experience operations for the world’s top support leaders. With 500+ customers, including Optum, Corpay, Sony, SharkNinja, and Allianz, we transform self-service, surface automation opportunities, and turn every agent into an expert.\nThe Role\nWe’re hiring a Senior DevOps / Platform Reliability Engineer to own the platform that powers our agentic CX product. You’ll build the CI/CD, infrastructure, and observability backbone that enables us to ship multi-agent systems safely to enterprise customers.\nIf you want to operate a production AI platform and use AI to help operate it, this role is for you.\nIn this role, you will collaborate with development, operations, and infrastructure teams to automate and streamline processes, build and maintain tools for deployment, monitoring, and operations, and troubleshoot issues across development and production environments.\nWhat You'll Do\nOwn and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.\nAutomate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.\nOperate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.\nManage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.\nOperate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.\nBuild and maintain Lambda workloads where event-driven or serverless architectures are the right fit.\nBuild observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.\nStrengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.\nDrive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.\nBuild and evolve our AI-native DevOps capabilities (see section below).\nPartner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.\nMonitor system performance and ensure reliability, scalability, and security across infrastructure and services.\nCollaborate with software engineering teams to support continuous integration and continuous delivery best practices.\nDocument infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.\nAgentic AI in DevOps\nYou’ll help define how Zingtree uses agentic AI to operate and improve our platform using modern AI operational practices.\nResponsibilities include:\nDesign and operate auto-remediation agents for common production toil such as certificate rotation, noisy pods, infrastructure drift, and flaky CI pipelines, with human-in-the-loop (HITL) controls for any destructive or customer-impacting actions.\nUse LLMs for incident triage and root cause analysis, including log and trace summarization, signal correlation, and first-draft postmortems that are always reviewed by humans.\nConnect AI agents to internal systems through the Model Context Protocol (MCP), including GitHub, Jira, PagerDuty, AWS, Kubernetes, Terraform, and related platforms, using scoped credentials, audit logging, and allow-listed access.\nApply AI-driven observability techniques, including anomaly detection on metrics, LLM-based log clustering, and alert deduplication and summarization on top of Prometheus and OpenTelemetry.\nEstablish operational guardrails such as prompt/version pinning, evaluation frameworks for agent behavior, cost and rate-limit controls, policy-as-code (OPA/Conftest) for AI-generated infrastructure changes, and clearly defined blast-radius controls.\nDefine best practices for AI coding assistants such as GitHub Copilot, Claude, and Amazon Q in infrastructure repositories, including review workflows, prompt design, and restrictions on auto-merged changes.\nTreat AI components as production systems with SLOs, observability, on-call readiness, runbooks, and rollback strategies for agents and prompts.\nAbout You\nRequired Qualifications\n5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.\nStrong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.\nHands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.\nStrong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.\nDeep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.\nExperience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.\nStrong observability experience using Prometheus, Grafana, and OpenTelemetry.\nExperience operating Argo CD at scale.\nExperience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.\nExperience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.\nExperience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.\nExperience with Lambda and event-driven architectures.\nComfortable working with Python, Bash, and Linux systems.\nStrong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.\nFamiliarity with vulnerability scanning and compliance tooling.\nNice to Have\nExperience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.\nExperience building or integrating MCP servers or deploying agent frameworks such as LangGraph or CrewAI in production environments.\nHow We Work\nWe bias toward automation over toil. If you do it twice, script it. If it pages twice, fix it.\nWe’re a small team with high ownership. You’ll help define standards, not just follow them.\nHumans stay in the loop for anything risky. AI accelerates decision-making but does not replace judgment.\nWe value blameless incident reviews, documented decisions, and fast feedback loops.\nWhat We Offer\nCompetitive compensation packages\nComprehensive health benefits:\n100% of employee premiums covered\n75%–80% of dependent premiums covered for most health, dental, and vision plans\n401(k) plans to support retirement planning (no employer matching currently)\nPaid parental leave\nUnlimited PTO\nFlexible remote work from anywhere\nUp to $200/month co-working reimbursement\nHome office stipend:\nUp to $500 for home office setup\n$100/month for internet, phone, and related expenses\nZingtree Values\nLead with Action\nWe are doers. We move quickly with purpose, take smart risks, learn fast, and focus on outcomes that benefit our customers and the business.\nPeople Really Matter\nWe win as a team. We care deeply about our customers and employees, helping each other achieve professional growth and meaningful impact.\nOwnership Leads to Results\nWhen we commit, we deliver. We operate with integrity, accountability, and high standards.\nExpertise Creates Value\nWe are learners. We continuously grow our knowledge, share expertise, and apply it to create meaningful results.\nTransparency Builds Trust\nWe communicate openly, honestly, and respectfully. We share information that matters and build trusted relationships through clarity and empathy.\nWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.","company":"Zingtree","rawCompany":"zingtree","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-04T17:04:23.831Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior DevOps / Platform Reliability Engineer","description":"About Zingtree\nZingtree is the next-generation intelligent process automation platform reimagining customer experience operations for the world’s top support leaders. With 500+ customers, including Optum, Corpay, Sony, SharkNinja, and Allianz, we transform self-service, surface automation opportunities, and turn every agent into an expert.\nThe Role\nWe’re hiring a Senior DevOps / Platform Reliability Engineer to own the platform that powers our agentic CX product. You’ll build the CI/CD, infrastructure, and observability backbone that enables us to ship multi-agent systems safely to enterprise customers.\nIf you want to operate a production AI platform and use AI to help operate it, this role is for you.\nIn this role, you will collaborate with development, operations, and infrastructure teams to automate and streamline processes, build and maintain tools for deployment, monitoring, and operations, and troubleshoot issues across development and production environments.\nWhat You'll Do\nOwn and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.\nAutomate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.\nOperate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.\nManage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.\nOperate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.\nBuild and maintain Lambda workloads where event-driven or serverless architectures are the right fit.\nBuild observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.\nStrengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.\nDrive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.\nBuild and evolve our AI-native DevOps capabilities (see section below).\nPartner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.\nMonitor system performance and ensure reliability, scalability, and security across infrastructure and services.\nCollaborate with software engineering teams to support continuous integration and continuous delivery best practices.\nDocument infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.\nAgentic AI in DevOps\nYou’ll help define how Zingtree uses agentic AI to operate and improve our platform using modern AI operational practices.\nResponsibilities include:\nDesign and operate auto-remediation agents for common production toil such as certificate rotation, noisy pods, infrastructure drift, and flaky CI pipelines, with human-in-the-loop (HITL) controls for any destructive or customer-impacting actions.\nUse LLMs for incident triage and root cause analysis, including log and trace summarization, signal correlation, and first-draft postmortems that are always reviewed by humans.\nConnect AI agents to internal systems through the Model Context Protocol (MCP), including GitHub, Jira, PagerDuty, AWS, Kubernetes, Terraform, and related platforms, using scoped credentials, audit logging, and allow-listed access.\nApply AI-driven observability techniques, including anomaly detection on metrics, LLM-based log clustering, and alert deduplication and summarization on top of Prometheus and OpenTelemetry.\nEstablish operational guardrails such as prompt/version pinning, evaluation frameworks for agent behavior, cost and rate-limit controls, policy-as-code (OPA/Conftest) for AI-generated infrastructure changes, and clearly defined blast-radius controls.\nDefine best practices for AI coding assistants such as GitHub Copilot, Claude, and Amazon Q in infrastructure repositories, including review workflows, prompt design, and restrictions on auto-merged changes.\nTreat AI components as production systems with SLOs, observability, on-call readiness, runbooks, and rollback strategies for agents and prompts.\nAbout You\nRequired Qualifications\n5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.\nStrong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.\nHands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.\nStrong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.\nDeep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.\nExperience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.\nStrong observability experience using Prometheus, Grafana, and OpenTelemetry.\nExperience operating Argo CD at scale.\nExperience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.\nExperience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.\nExperience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.\nExperience with Lambda and event-driven architectures.\nComfortable working with Python, Bash, and Linux systems.\nStrong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.\nFamiliarity with vulnerability scanning and compliance tooling.\nNice to Have\nExperience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.\nExperience building or integrating MCP servers or deploying agent frameworks such as LangGraph or CrewAI in production environments.\nHow We Work\nWe bias toward automation over toil. If you do it twice, script it. If it pages twice, fix it.\nWe’re a small team with high ownership. You’ll help define standards, not just follow them.\nHumans stay in the loop for anything risky. AI accelerates decision-making but does not replace judgment.\nWe value blameless incident reviews, documented decisions, and fast feedback loops.\nWhat We Offer\nCompetitive compensation packages\nComprehensive health benefits:\n100% of employee premiums covered\n75%–80% of dependent premiums covered for most health, dental, and vision plans\n401(k) plans to support retirement planning (no employer matching currently)\nPaid parental leave\nUnlimited PTO\nFlexible remote work from anywhere\nUp to $200/month co-working reimbursement\nHome office stipend:\nUp to $500 for home office setup\n$100/month for internet, phone, and related expenses\nZingtree Values\nLead with Action\nWe are doers. We move quickly with purpose, take smart risks, learn fast, and focus on outcomes that benefit our customers and the business.\nPeople Really Matter\nWe win as a team. We care deeply about our customers and employees, helping each other achieve professional growth and meaningful impact.\nOwnership Leads to Results\nWhen we commit, we deliver. We operate with integrity, accountability, and high standards.\nExpertise Creates Value\nWe are learners. We continuously grow our knowledge, share expertise, and apply it to create meaningful results.\nTransparency Builds Trust\nWe communicate openly, honestly, and respectfully. We share information that matters and build trusted relationships through clarity and empathy.\nWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.","datePosted":"2026-08-04T17:04:23.831Z","dateModified":"2026-08-04T17:04:23.831Z","hiringOrganization":{"@type":"Organization","name":"Zingtree","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"76bfcd7b8a26000a7070ab19"},"url":"https://jobsearcher.com/jobs/76bfcd7b8a26000a7070ab19"}}