Jobs and Careers
SE

Senior Staff, Product Manager- AI Evaluation & Quality

ServiceNow
Santa Clara, United Statesfull_timeVerifiedPosted 22 Jul 2026

About the role

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

Job Description

About the team

This role sits within APEX (AI Foundations), the team building the platform-level infrastructure behind ServiceNow's enterprise AI strategy. You'll work alongside engineering, data, and applied research partners to make evaluation a durable, compounding advantage.

About the role

Agentic AI is only as trustworthy as the evaluation discipline behind it. ServiceNow is building the evaluation methodology, tooling, and closed-loop production system that will define how good is good enough for AI specialists and multi-agent frameworks across the enterprise — and we're looking for the product leader to help own that discipline. This is a rare chance to build a category-defining evaluation platform from the ground up, at a company where the outcome shapes how every business unit ships AI, not just one team's roadmap.

Why this role matters

Evaluation is the foundational differentiator for enterprise agentic AI. As frontier models commoditize, defensible advantage shifts to the orchestration and quality layer — how reliably an AI specialist performs in a customer's production environment. This role sits at exactly that inflection point:

  • You set direction in undefined space — there is no existing playbook to inherit, and the discipline you build becomes the standard others build on.
  • Your impact is horizontal and enterprise-wide by design, driven through influence, tooling, and a Center of Excellence rather than headcount.
  • You'll be the definitive point of reference for AI specialist quality across the portfolio — deep technical and methodological ownership without the dilution of people management.
  • Success is visible and concrete: your framework adopted across business units, a closed-loop system demonstrably improving specialists release-over-release, and tooling used self-serve by teams you never directly staffed.

The impact you'll make

  • Evaluation strategy and framework: Define the end-to-end evaluation methodology across ServiceNow's AI specialist portfolio — golden datasets, LLM-as-judge calibration, failure taxonomy, and a multi-layer metric model spanning agent behavior, user experience, and business impact.
  • Closed-loop evaluation platform: Move evaluation from a one-time release gate to a continuous improvement engine, where production telemetry feeds failure analysis, targeted evaluation expansion, and redeployment — making live signal, not synthetic testing, the primary driver of quality.
  • Cross-functional influence at scale: Partner with a federated Center of Excellence model — central methodology and tooling, embedded practice in each business unit — acting as the internal consulting function that helps teams stand up evaluation without rebuilding infrastructure from scratch.
  • Platform and tooling: Drive evaluation from bespoke effort to reusable, self-serve platform capability — evaluation infrastructure, data tooling, and calibration systems that make closed-loop evaluation available to every team building AI specialists.

What Success Looks Like

  • Your evaluation framework is adopted across multiple business units.
  • The closed-loop platform demonstrably drives specialist improvement release-over-release, sourced from real production signal.
  • Evaluation tooling is used self-serve by teams the Center of Excellence never directly staffed.
  • Your quality bar becomes the internally recognized authoritative standard for AI specialist readiness.

Qualifications

  • 12+ years of software product management experience
  • Deep experience in AI/ML product management, ideally with hands-on exposure to LLM evaluation, agentic systems, or applied ML quality frameworks.
  • A track record of defining methodology or standards in ambiguous, cross-team problem spaces — not just executing an existing roadmap.
  • Strong technical fluency — comfortable in the details of evaluation pipelines, data platforms, and production telemetry, not just the product narrative around them.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

ServiceNow

View company profile →