Senior Manager, Site Reliability and Production Delivery Management (US)
TDAbout the role
Work Location:
Mount Laurel, New Jersey, United States of AmericaHours:
40Pay Details:
$143,720 - $223,080 USDTD is committed to providing fair and equitable compensation opportunities to all colleagues. Growth opportunities and skill development are defining features of the colleague experience at TD. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The base pay actually offered may vary based upon the candidate's skills and experience, job-related knowledge, geographic location, and other specific business and organizational needs.
As a candidate, you are encouraged to ask compensation related questions and have an open dialogue with your recruiter who can provide you more specific details for this role.
Line of Business:
Technology SolutionsJob Description:
Seeking a highly experienced Senior Manager, Site Reliability and Production Delivery Management to lead reliability, availability, production deployment, and operational excellence practices across critical technology platforms. This leader will be responsible for defining and executing the strategy for observability, SRE practices, release governance, change management, and production deployment processes to ensure highly available, secure, scalable, and resilient customer-facing applications and services.
The role will partner closely with Application Development, Infrastructure, Platform Engineering, Architecture, Operations, and Product Management teams to drive a culture of reliability, automation, continuous improvement, and operational accountability. The successful candidate will establish enterprise standards for monitoring, incident management, release engineering, deployment automation, and service reliability while enabling faster and safer delivery of business capabilities. This leader will drive operational excellence, resiliency, automation, and governance while enabling the safe and efficient delivery of technology solutions.
Key Responsibilities
- Lead and mature platform Observability and SRE practices, including monitoring, alerting, SLOs, SLIs, error budgets, and operational readiness.
- Establish standards for dashboards, logging, tracing, incident response, runbooks, and on-call support.
- Drive continuous improvement in reliability, resiliency, automation, and operational efficiency.
- Lead release planning, deployment governance, and change management processes across multiple delivery teams.
- Oversee production deployments, deployment readiness reviews, risk management, and rollback strategies.
- Partner with Engineering, Architecture, Infrastructure, Operations, and Product teams to improve service stability and delivery performance.
- Provide leadership during major incidents, post-incident reviews, and service restoration activities.
- Track and report key operational metrics, including availability, MTTR, deployment success rate, and change failure rate.
Leadership & Strategy
- Define and execute the platform roadmap for Observability, SRE, Release Management, and Production Deployment.
- Build and lead high-performing teams responsible for reliability engineering, monitoring, release coordination, and deployment governance.
- Establish service reliability objectives, operational standards, and performance metrics aligned with business goals.
- Champion a reliability-first culture across technology delivery and operations organizations.
- Provide executive-level reporting on service health, availability, deployment performance, and operational risks.
Site Reliability Engineering (SRE)
- Develop and mature SRE practices, including SLOs, SLIs, error budgets, capacity planning, resiliency engineering, and operational readiness reviews.
- Drive adoption of automation to reduce operational toil and improve service stability.
- Lead incident management, post-incident reviews, root cause analysis, and continuous improvement initiatives.
- Ensure application teams incorporate reliability, recoverability, and non-functional requirements throughout the software development lifecycle.
- Establish standards for runbooks, operational playbooks, escalation procedures, and on-call readiness.
Observability & Monitoring
- Establish platform observability standards covering metrics, logs, traces, business KPIs, and customer experience monitoring.
- Drive implementation of dashboards that provide executive-level, operational, and application-level visibility.
- Ensure monitoring solutions provide actionable insights and proactive detection of service degradation.
- Define alerting strateg
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s