Staff Software Engineer, Forecasting
Zeta GlobalAbout the role
WHO WE ARE
Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform – powered by one of the industry’s largest proprietary databases and AI. Our enterprise customers across multiple verticals are empowered to personalize experiences with consumers at an individual level across every channel, delivering better results for marketing programs. Zeta was founded in 2007 by David A. Steinberg and John Sculley and is headquartered in New York City with offices around the world. To learn more, go to www.zetaglobal.com.
THE ROLE
We are hiring a hands-on Staff Software Engineer to provide technical leadership for our Forecasting and Recommendations platforms, with a strong focus on production-grade AI and agentic systems.
This role centers on designing, building, and operating high-throughput, low-latency distributed systems that power forecasting, recommendations, and AI-driven decisioning at scale. You will work deeply in backend systems, infrastructure, and AI application architecture, while remaining accountable for reliability, observability, and operational excellence.
We are looking for an engineer who can contribute immediately, has shouldered real production incidents, and brings strong judgment around building stable, observable, and scalable systems, including modern agentic and LLM-powered applications.
WHAT YOU'LL DO
- Design, build, and operate systems supporting forecasting, recommendations, and agentic AI workflows in production.
- Write production-quality code daily; own services end-to-end from design through on-call and incident resolution.
- Architect low-latency, high-throughput SaaS services, including APIs, data pipelines, model inference, and agent orchestration.
- Build and maintain production-grade agentic applications, including tool-using agents, workflow orchestration, and guardrails.
- Work fluently with foundational LLMs (e.g., GPT, Claude, Gemini Pro), selecting appropriate models and deployment patterns based on latency, cost, and reliability tradeoffs.
- Use frameworks and tooling such as LangChain, voice agents, and related ecosystems to accelerate development—while enforcing production discipline.
- Embrace AI-assisted development workflows (e.g., Cursor, GitHub Copilot, vibe coding paradigms) to move quickly without sacrificing quality.
- Champion observability and reliability: metrics, logging, tracing, alerting, and post-incident analysis.
- Lead and participate in production incident response, retrospectives, and systemic fixes.
- Identify architectural risks early and make design decisions that prevent outages and scalability issues.
- Reduce complexity across services, infrastructure, and processes to improve stability and team velocity.
- Provide technical guidance across teams and participate in architectural reviews beyond your immediate domain.
WHAT WE'RE LOOKING FOR
- 10+ years of professional software engineering experience building and operating production-grade distributed systems.
- A strong track record of hands-on ownership of business-critical services, including measurable improvements in latency, throughput, stability, or cost.
- Deep expertise in systems design, including service boundaries, concurrency, data modeling, failure handling, and scalability tradeoffs.
- Production experience supporting machine learning–driven systems
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s