{"schemaVersion":"jobsearcher.job.v1","id":"91d02867ec2dfa92239b1ca0","url":"https://jobsearcher.com/jobs/91d02867ec2dfa92239b1ca0","canonicalUrl":"https://jobsearcher.com/jobs/91d02867ec2dfa92239b1ca0","title":"Principal DevOps Engineer","description":"WHO WE ARE\nZeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform – powered by one of the industry’s largest proprietary databases and AI. Our enterprise customers across multiple verticals are empowered to personalize experiences with consumers at an individual level across every channel, delivering better results for marketing programs. Zeta was founded in 2007 by David A. Steinberg and John Sculley and is headquartered in New York City with offices around the world. To learn more, go to www.zetaglobal.com.\nRole Overview\nWe are seeking a Principal DevOps Engineer to serve as a transformative force in how ZetaGlobal builds, deploys, and operates software at scale. This is not a maintenance role. You will be a DevOps disruptor: someone who challenges the status quo, reimagines deployment pipelines, and empowers hundreds of developers across multiple teams to ship code to production safely, multiple times per day, concurrently.\nYour prime responsibility is to enable true continuous integration and continuous deployment (CI/CD) directly to production using canary releases, blue/green deployments, incremental rollout pipelines, feature flag-driven releases, and any other proven strategy that delivers speed with safety. You will architect and operate these systems within a regulated, globally compliant environment spanning GDPR, CCPA, and SOC 2 requirements.\nIn addition, you will serve as a Site Reliability Engineer (SRE) leader, ensuring safe operations, incident readiness, and platform stability as we continue to scale. You will influence both DevOps/SRE practices and software architecture decisions, simplifying and streamlining operational management across the organization.\nKey Responsibilities\nCI/CD & Deployment Excellence\nDesign, build, and operate production-grade CI/CD pipelines enabling multiple developers on multiple teams to deploy concurrently to production, multiple times daily, with zero-downtime guarantees.\nImplement and optimize advanced deployment strategies including canary releases, blue/green deployments, rolling updates, incremental rollouts, and feature flag-gated releases via Statsig.\nBuild self-service deployment tooling that empowers developers to own their release process while enforcing safety guardrails, automated rollback triggers, and automate compliance gates.\nEstablish deployment observability with real-time canary analysis, automated health scoring, and progressive delivery metrics integrated with Grafana, Prometheus, and Honeycomb.\nChampion CI/CD workflows using GitLab CI/CD, Helm charts, and Terraform to ensure infrastructure and application deployments are version-controlled, auditable, and reproducible.\nPlatform Reliability & SRE\nDefine and enforce SLOs/SLIs/SLAs across services, establishing error budgets that balance velocity with reliability.\nLead incident response processes, including on-call rotations, runbook development, blameless postmortems, and incident command structure.\nDesign and implement robust observability stacks leveraging Grafana, Prometheus, Loki, and Honeycomb for metrics, logging, tracing, and alerting at scale.\nProactively identify and eliminate reliability risks through chaos engineering, load testing, capacity planning, and failure mode analysis.\nReduce operational toil through automation, self-healing infrastructure patterns, and intelligent alerting to minimize mean time to detection (MTTD) and recovery (MTTR).\nInfrastructure & Architecture\nManage and optimize AWS infrastructure spanning EC2, SQS, DynamoDB, and related services with Infrastructure as Code (Terraform) best practices.\nDesign and operate Kafka-based event streaming infrastructure for high-throughput, low-latency data pipelines supporting real-time marketing and analytics workloads.\nEnsure robust networking across the platform, including DNS management, service mesh configuration, load balancing, TCP/IP optimization, routing policies, and VPC architecture.\nManage containerization strategy using Docker, ensuring efficient image builds, vulnerability scanning, registry management, and runtime security.\nSupport data infrastructure operations across Snowflake, MySQL, and other database platforms, collaborating with data engineering teams on reliability and performance.\nCompliance, Security & Governance\nEmbed compliance controls directly into CI/CD pipelines, ensuring automated enforcement of GDPR, CCPA, and SOC 2 requirements at every stage of the software delivery lifecycle.\nImplement audit trails, change management controls, and deployment approval workflows required by regulatory frameworks in the MarTech and AdTech domains.\nCollaborate with Security and Legal teams to ensure infrastructure and deployment processes meet global compliance obligations across all operating regions.\nMaintain awareness of evolving privacy regulations (ePrivacy, state-level US laws, international data residency requirements) and proactively adapt infrastructure accordingly.\nTechnical Leadership & Influence\nServe as a technical leader and DevOps disruptor, challenging legacy processes and introducing modern practices that dramatically improve developer velocity and operational safety.\nInfluence software architecture decisions to simplify and streamline operational management, advocating for patterns that are deployment-friendly, observable, and resilient by design.\nClearly communicate complex technical strategies to engineering leadership, product stakeholders, and cross-functional teams to build alignment and drive adoption.\nDevelop reference architectures, internal standards, and golden path templates that codify best practices and accelerate onboarding of new services and teams.\nParticipate in on-call rotations and lead by example in incident response, demonstrating the operational discipline expected across the engineering organization.\nRequired Qualifications\n10+ years of progressive experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering roles, with demonstrated impact at staff or principal level.\nExpert-level Kubernetes knowledge, including cluster administration, Helm chart authoring, custom controllers/operators, network policies, RBAC, and multi-cluster management on AWS EKS.\nDeep expertise in CI/CD pipeline architecture and advanced deployment strategies (canary, blue/green, progressive delivery, feature flag integration) at scale.\nStrong proficiency with Infrastructure as Code using Terraform, including module design, state management, and multi-environment orchestration.\nExpert knowledge of Docker containerization, including multi-stage builds, security hardening, image optimization, and container runtime management.\nProduction experience with Apache Kafka, including cluster management, topic design, consumer group strategies, and operational monitoring for high-throughput streaming workloads.\nStrong networking fundamentals: DNS (Route 53, internal DNS), TCP/IP, routing, API Gateway, load balancing (ALB/NLB), service mesh, VPC peering, transit gateways, and network troubleshooting.\nExtensive AWS experience spanning EKS, EC2, SQS, DynamoDB, IAM, VPC, CloudWatch, and related services in production environments.\nHands-on experience with observability platforms: Grafana (dashboards, alerting), Prometheus (metrics, PromQL), Loki (log aggregation), and Honeycomb (distributed tracing, BubbleUp analysis).\nWorking familiarity with multiple language stacks including Node.js, React, Python, Java, and Ruby, sufficient to understand build systems, dependency management, and runtime characteristics.\nExperience operating within regulated environments, with practical knowledge of GDPR, CCPA, SOC 2, and compliance automation in MarTech or AdTech domains.\nProven ability to influence engineering culture, drive adoption of new practices, and communicate complex technical strategies clearly to both technical and non-technical stakeholders.\nDemonstrated experience with GitLab CI/CD pipelines, including advanced pipeline features such as parent-child pipelines, dynamic environments, and security scanning integration.\nPreferred Qualifications\nAWS certifications: Solutions Architect Professional, DevOps Engineer Professional, or Security Specialty.\nExperience with Statsig or similar feature flag and experimentation platforms for progressive delivery and A/B testing in production.\nBackground in chaos engineering tools and practices (Gremlin, Litmus, Chaos Monkey) for proactive resilience validation.\nExperience building internal developer platforms (IDPs) or platform-as-a-product organizations.\nFamiliarity with FinOps practices and cloud cost optimization strategies.\nContributions to open-source DevOps/SRE tools or active participation in the broader infrastructure community.\nExperience with service mesh technologies (Istio, Linkerd) for advanced traffic management and security.\nBENEFITS & PERKS\nUnlimited PTO\nExcellent medical, dental, and vision coverage\nEmployee Equity\nEmployee Discounts, Virtual Wellness Classes, and Pet Insurance And more!!\nSALARY RANGE\nThe salary range for this role is $180,000 - $210,000, depending on location and experience.\nPEOPLE & CULTURE AT ZETA\nZeta considers applicants for employment without regard to, and does not discriminate on the basis of an individual’s sex, race, color, religion, age, disability, status as a veteran, or national or ethnic origin; nor does Zeta discriminate on the basis of sexual orientation, gender identity or expression.\nWe’re committed to building a workplace culture of trust and belonging, so everyone feels invited to bring their whole selves to work. We provide a forum for employees to celebrate, support and advocate for one another. Learn more about our commitment to diversity, equity and inclusion here: https://zetaglobal.com/blog/a-look-into-zetas-ergs/\nZETA IN THE NEWS!\nhttps://zetaglobal.com/press/?cat=press-releases\n\n#LI-YW1","company":"Zetaglobal","rawCompany":"zetaglobal","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-07-31T11:29:02.146Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal DevOps Engineer","description":"WHO WE ARE\nZeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform – powered by one of the industry’s largest proprietary databases and AI. Our enterprise customers across multiple verticals are empowered to personalize experiences with consumers at an individual level across every channel, delivering better results for marketing programs. Zeta was founded in 2007 by David A. Steinberg and John Sculley and is headquartered in New York City with offices around the world. To learn more, go to www.zetaglobal.com.\nRole Overview\nWe are seeking a Principal DevOps Engineer to serve as a transformative force in how ZetaGlobal builds, deploys, and operates software at scale. This is not a maintenance role. You will be a DevOps disruptor: someone who challenges the status quo, reimagines deployment pipelines, and empowers hundreds of developers across multiple teams to ship code to production safely, multiple times per day, concurrently.\nYour prime responsibility is to enable true continuous integration and continuous deployment (CI/CD) directly to production using canary releases, blue/green deployments, incremental rollout pipelines, feature flag-driven releases, and any other proven strategy that delivers speed with safety. You will architect and operate these systems within a regulated, globally compliant environment spanning GDPR, CCPA, and SOC 2 requirements.\nIn addition, you will serve as a Site Reliability Engineer (SRE) leader, ensuring safe operations, incident readiness, and platform stability as we continue to scale. You will influence both DevOps/SRE practices and software architecture decisions, simplifying and streamlining operational management across the organization.\nKey Responsibilities\nCI/CD & Deployment Excellence\nDesign, build, and operate production-grade CI/CD pipelines enabling multiple developers on multiple teams to deploy concurrently to production, multiple times daily, with zero-downtime guarantees.\nImplement and optimize advanced deployment strategies including canary releases, blue/green deployments, rolling updates, incremental rollouts, and feature flag-gated releases via Statsig.\nBuild self-service deployment tooling that empowers developers to own their release process while enforcing safety guardrails, automated rollback triggers, and automate compliance gates.\nEstablish deployment observability with real-time canary analysis, automated health scoring, and progressive delivery metrics integrated with Grafana, Prometheus, and Honeycomb.\nChampion CI/CD workflows using GitLab CI/CD, Helm charts, and Terraform to ensure infrastructure and application deployments are version-controlled, auditable, and reproducible.\nPlatform Reliability & SRE\nDefine and enforce SLOs/SLIs/SLAs across services, establishing error budgets that balance velocity with reliability.\nLead incident response processes, including on-call rotations, runbook development, blameless postmortems, and incident command structure.\nDesign and implement robust observability stacks leveraging Grafana, Prometheus, Loki, and Honeycomb for metrics, logging, tracing, and alerting at scale.\nProactively identify and eliminate reliability risks through chaos engineering, load testing, capacity planning, and failure mode analysis.\nReduce operational toil through automation, self-healing infrastructure patterns, and intelligent alerting to minimize mean time to detection (MTTD) and recovery (MTTR).\nInfrastructure & Architecture\nManage and optimize AWS infrastructure spanning EC2, SQS, DynamoDB, and related services with Infrastructure as Code (Terraform) best practices.\nDesign and operate Kafka-based event streaming infrastructure for high-throughput, low-latency data pipelines supporting real-time marketing and analytics workloads.\nEnsure robust networking across the platform, including DNS management, service mesh configuration, load balancing, TCP/IP optimization, routing policies, and VPC architecture.\nManage containerization strategy using Docker, ensuring efficient image builds, vulnerability scanning, registry management, and runtime security.\nSupport data infrastructure operations across Snowflake, MySQL, and other database platforms, collaborating with data engineering teams on reliability and performance.\nCompliance, Security & Governance\nEmbed compliance controls directly into CI/CD pipelines, ensuring automated enforcement of GDPR, CCPA, and SOC 2 requirements at every stage of the software delivery lifecycle.\nImplement audit trails, change management controls, and deployment approval workflows required by regulatory frameworks in the MarTech and AdTech domains.\nCollaborate with Security and Legal teams to ensure infrastructure and deployment processes meet global compliance obligations across all operating regions.\nMaintain awareness of evolving privacy regulations (ePrivacy, state-level US laws, international data residency requirements) and proactively adapt infrastructure accordingly.\nTechnical Leadership & Influence\nServe as a technical leader and DevOps disruptor, challenging legacy processes and introducing modern practices that dramatically improve developer velocity and operational safety.\nInfluence software architecture decisions to simplify and streamline operational management, advocating for patterns that are deployment-friendly, observable, and resilient by design.\nClearly communicate complex technical strategies to engineering leadership, product stakeholders, and cross-functional teams to build alignment and drive adoption.\nDevelop reference architectures, internal standards, and golden path templates that codify best practices and accelerate onboarding of new services and teams.\nParticipate in on-call rotations and lead by example in incident response, demonstrating the operational discipline expected across the engineering organization.\nRequired Qualifications\n10+ years of progressive experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering roles, with demonstrated impact at staff or principal level.\nExpert-level Kubernetes knowledge, including cluster administration, Helm chart authoring, custom controllers/operators, network policies, RBAC, and multi-cluster management on AWS EKS.\nDeep expertise in CI/CD pipeline architecture and advanced deployment strategies (canary, blue/green, progressive delivery, feature flag integration) at scale.\nStrong proficiency with Infrastructure as Code using Terraform, including module design, state management, and multi-environment orchestration.\nExpert knowledge of Docker containerization, including multi-stage builds, security hardening, image optimization, and container runtime management.\nProduction experience with Apache Kafka, including cluster management, topic design, consumer group strategies, and operational monitoring for high-throughput streaming workloads.\nStrong networking fundamentals: DNS (Route 53, internal DNS), TCP/IP, routing, API Gateway, load balancing (ALB/NLB), service mesh, VPC peering, transit gateways, and network troubleshooting.\nExtensive AWS experience spanning EKS, EC2, SQS, DynamoDB, IAM, VPC, CloudWatch, and related services in production environments.\nHands-on experience with observability platforms: Grafana (dashboards, alerting), Prometheus (metrics, PromQL), Loki (log aggregation), and Honeycomb (distributed tracing, BubbleUp analysis).\nWorking familiarity with multiple language stacks including Node.js, React, Python, Java, and Ruby, sufficient to understand build systems, dependency management, and runtime characteristics.\nExperience operating within regulated environments, with practical knowledge of GDPR, CCPA, SOC 2, and compliance automation in MarTech or AdTech domains.\nProven ability to influence engineering culture, drive adoption of new practices, and communicate complex technical strategies clearly to both technical and non-technical stakeholders.\nDemonstrated experience with GitLab CI/CD pipelines, including advanced pipeline features such as parent-child pipelines, dynamic environments, and security scanning integration.\nPreferred Qualifications\nAWS certifications: Solutions Architect Professional, DevOps Engineer Professional, or Security Specialty.\nExperience with Statsig or similar feature flag and experimentation platforms for progressive delivery and A/B testing in production.\nBackground in chaos engineering tools and practices (Gremlin, Litmus, Chaos Monkey) for proactive resilience validation.\nExperience building internal developer platforms (IDPs) or platform-as-a-product organizations.\nFamiliarity with FinOps practices and cloud cost optimization strategies.\nContributions to open-source DevOps/SRE tools or active participation in the broader infrastructure community.\nExperience with service mesh technologies (Istio, Linkerd) for advanced traffic management and security.\nBENEFITS & PERKS\nUnlimited PTO\nExcellent medical, dental, and vision coverage\nEmployee Equity\nEmployee Discounts, Virtual Wellness Classes, and Pet Insurance And more!!\nSALARY RANGE\nThe salary range for this role is $180,000 - $210,000, depending on location and experience.\nPEOPLE & CULTURE AT ZETA\nZeta considers applicants for employment without regard to, and does not discriminate on the basis of an individual’s sex, race, color, religion, age, disability, status as a veteran, or national or ethnic origin; nor does Zeta discriminate on the basis of sexual orientation, gender identity or expression.\nWe’re committed to building a workplace culture of trust and belonging, so everyone feels invited to bring their whole selves to work. We provide a forum for employees to celebrate, support and advocate for one another. Learn more about our commitment to diversity, equity and inclusion here: https://zetaglobal.com/blog/a-look-into-zetas-ergs/\nZETA IN THE NEWS!\nhttps://zetaglobal.com/press/?cat=press-releases\n\n#LI-YW1","datePosted":"2026-07-31T11:29:02.146Z","dateModified":"2026-07-31T11:29:02.146Z","hiringOrganization":{"@type":"Organization","name":"Zetaglobal","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"91d02867ec2dfa92239b1ca0"},"url":"https://jobsearcher.com/jobs/91d02867ec2dfa92239b1ca0"}}