Technical Product Manager II, Site Reliability Engineering
The New York TimesAbout the role
<div class="content-intro"><div id="labeledImage.LOCATION" class="WIFG" data-automation-id="decorationWrapper"> <div class="WNHJ"> <div id="labeledImage.LOCATION--uid38" class="WE-Y WMXY WBAB WF0Y" data-automation-id="responsiveMonikerInput" data-metadata-id="labeledImage.LOCATION" data-uxi-form-item-child-list-index="0"> <div class="WJ-Y"> <p><strong>The <a href="https://www.nytco.com/company/mission-and-values/" target="_blank"><u>mission</u></a> of The New York Times is to seek the truth and help people understand the world. That means independent journalism is at the heart of all we do as a company. It’s why we have a world-renowned newsroom that sends journalists to report on the ground from nearly 160 countries. It’s why we focus deeply on how our readers will experience our journalism, from print to audio to a world-class digital and app destination. And it’s why our business strategy centers on making journalism so good that it’s worth paying for. </strong></p> </div> </div> </div> </div></div><p><strong>Mission Overview & Responsibilities</strong><span style="font-weight: 400;">: </span></p> <p data-pm-slice="1 1 []">At The New York Times, our Site Reliability Engineering (SRE) team is central to how we design, test, and operate the systems that support our most critical customer experiences. We're looking for a Technical Product Manager to lead the strategy for reliability programs and platforms that help teams ship resilient systems with confidence. These programs include operational readiness, load and chaos testing, observability, and incident readiness.</p> <p>You'll partner with SRE, platform infrastructure, and product engineering teams to define the standards, tooling, and practices that improve operational readiness across hundreds of services. You will focus on building scalable reliability programs and experiences—not running cloud infrastructure or managing operational tickets.</p> <p>You will be the product lead for a portfolio of SRE programs that includes:</p> <ul> <li> <p>Operational Readiness and Always Ready – reliability models, scorecards, production readiness reviews, and reliability signals that support operating reviews and critical customer journeys.</p> </li> <li> <p>Load, Chaos & Disaster Recovery Testing – platforms and practices that validate how systems behave under high traffic, zonal or regional failures, and degraded conditions.</p> </li> <li> <p>Observability & Incident Readiness – opinionated defaults for metrics, logs, traces, dashboards, alerts, and runbooks that i
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s