Principal DevOps Engineer
Zeta GlobalAbout the role
WHO WE ARE
Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform – powered by one of the industry’s largest proprietary databases and AI. Our enterprise customers across multiple verticals are empowered to personalize experiences with consumers at an individual level across every channel, delivering better results for marketing programs. Zeta was founded in 2007 by David A. Steinberg and John Sculley and is headquartered in New York City with offices around the world. To learn more, go to www.zetaglobal.com.
Role Overview
We are seeking a Principal DevOps Engineer to serve as a transformative force in how ZetaGlobal builds, deploys, and operates software at scale. This is not a maintenance role. You will be a DevOps disruptor: someone who challenges the status quo, reimagines deployment pipelines, and empowers hundreds of developers across multiple teams to ship code to production safely, multiple times per day, concurrently.
Your prime responsibility is to enable true continuous integration and continuous deployment (CI/CD) directly to production using canary releases, blue/green deployments, incremental rollout pipelines, feature flag-driven releases, and any other proven strategy that delivers speed with safety. You will architect and operate these systems within a regulated, globally compliant environment spanning GDPR, CCPA, and SOC 2 requirements.
In addition, you will serve as a Site Reliability Engineer (SRE) leader, ensuring safe operations, incident readiness, and platform stability as we continue to scale. You will influence both DevOps/SRE practices and software architecture decisions, simplifying and streamlining operational management across the organization.
Key Responsibilities
CI/CD & Deployment Excellence
- Design, build, and operate production-grade CI/CD pipelines enabling multiple developers on multiple teams to deploy concurrently to production, multiple times daily, with zero-downtime guarantees.
- Implement and optimize advanced deployment strategies including canary releases, blue/green deployments, rolling updates, incremental rollouts, and feature flag-gated releases via Statsig.
- Build self-service deployment tooling that empowers developers to own their release process while enforcing safety guardrails, automated rollback triggers, and automate compliance gates.
- Establish deployment observability with real-time canary analysis, automated health scoring, and progressive delivery metrics integrated with Grafana, Prometheus, and Honeycomb.
- Champion CI/CD workflows using GitLab CI/CD, Helm charts, and Terraform to ensure infrastructure and application deployments are version-controlled, auditable, and reproducible.
Platform Reliability & SRE
- Define and enforce SLOs/SLIs/SLAs across services, establishing error budgets that balance velocity with reliability.
- Lead incident response processes, including on-call rotations, runbook development, blameless postmortems, and incident command structure.
- Design and implement robust observability stacks leveraging Grafana, Prometheus, Loki, and Honeycomb for metrics, logging, tracing, and alerting at scale.
- Proactively identify and eliminate reliability risks through chaos engineering, load testing, capacity planning, and failure mode analysis.
- Reduce operational toil through automation, self-healing infrastructure patterns, and intelligent alerting to minimize mean time to detection (MTTD) and recovery (MTTR).
Infrastructure & Architecture
- Manage and optimize AWS infrastructure spanning EC2, SQS, DynamoDB, and related services with Infrastructure as Code (Terraform) best practices.
- Design and operate Kafka-based event streaming infrastructure for high-throughput, low-latency data pipelines supporting real-time marketing and analytics workloads.
- Ensure robust networking across the platform, including DNS management, service mesh configuration, load balancing, TCP/IP optimization, routing policies, and VPC architecture.
- Manage containerization strategy using Docker, ensuring efficient image builds, vulnerability scanning, registry management, and runtime security.
- Support data infrastructure operations across Snowflake, MySQL, and other database platforms, collaborating with data engineering teams on reliability and pe
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s