Jobs and Careers
ZI

Senior Manager, Incident Management

Zillow
Remote-USA, United States, United StatesRemotefull_timeVerifiedPosted 7 Aug 2026
💰 $211,600/yr($132,400/yr$211,600/yr)

About the role

About the team

Zillow Group is seeking a Senior Manager, Incident Management (M4) to lead our incident and problem management function and elevate how we drive operational excellence across production systems. This role owns the strategy, process, and people behind incident response and problem management, ensuring we not only resolve incidents quickly but prevent recurrence and continuously raise the bar on reliability.

At the M4 level, success is anchored in four core areas: Program Leadership: building and scaling the incident and problem management practice across the organization; Problem Management: driving structured root cause analysis and systemic fixes that reduce repeat incidents; AI-Enabled Operations: leveraging AI workflows and tooling to increase the speed, quality, and actionability of incident and problem outputs for internal partners; and People Leadership: developing a high-performing team of incident managers and setting the standard for operational rigor.

About the role

Responsibilities

Incident Management Leadership

  • Own the end-to-end incident management program, including process design, tooling, and governance across the organization

  • Serve as executive escalation point and senior decision-maker for critical, high-severity, or cross-functional incidents

  • Set and enforce standards for incident severity classification, escalation paths, and communication protocols

  • Partner with Engineering, Product, and business leadership to align incident response with business priorities and risk tolerance

  • Drive executive-level incident communications, ensuring leadership has clear, timely, and accurate visibility into impact and status

Problem Management

  • Build and own a formal problem management practice that connects incident trends to systemic root causes

  • Ensure every significant incident produces a rigorous, blameless root cause analysis (RCA) with clearly owned, tracked corrective actions

  • Establish mechanisms to identify recurring issues, chronic risks, and process gaps across incident history

  • Hold cross-functional partners accountable for closing problem records and remediation items on committed timelines

  • Report on problem management outcomes and reliability trends to leadership, tying them to measurable risk reduction

Leveraging AI Workflows for Speed & Quality

  • Champion the adoption of AI-powered tooling and workflows across incident detection, triage, summarization, and RCA drafting

  • Design and continuously improve AI-assisted workflows that turn raw incident and problem data into clear, actionable insights for internal partners

  • Ensure AI-generated call-outs, summaries, and reports meet a high bar for accuracy, relevance, and actionability before reaching stakeholders

  • Identify opportunities to automate repetitive operational tasks (documentation, status updates, trend analysis) to free the team to focus on higher-value judgment work

  • Partner with Engineering and Data teams to pilot, evaluate, and scale new AI capabilities that improve mean-time-to-resolution and mean-time-to-detection

People & Team Leadership

  • Hire, coach, and develop a team of incident managers, building depth and bench strength across severity levels

  • Set clear performance expectations and career growth paths for the team

  • Establish on-call structures, workload balance, and rotations that sustain team health and reliability coverage

  • Foster a culture of ownership, continuous improvement, and blameless learning within the team

Operational Excellence & Continuous Improvement

  • Define and track key metrics (MTTR, MTTD, recurrence rate, action-item closure rate) to measure program health and impact

  • Continuously refine runbooks, tooling, and workflows based on retrospectives and data trends

  • Facilitate post-incident reviews for major incidents, ensuring lessons learned translate into concrete process or system changes

  • Benchmark practices against industry standards and bring in outside best practices where relevant

Scope & Impact

  • Owns the incident and problem management strategy and roadmap for the

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Zillow

View company profile →