Senior Manager, Site Reliability Engineering
Wells FargoAbout the role
About this role:
Wells Fargo is seeking a Systems Operations Senior Manager, Site Reliability Engineering (SRE) to lead a team of engineers responsible for the reliability, availability, performance, scalability, and operational excellence of critical customer-facing and enterprise technology platforms. This leader drives resiliency engineering, observability, incident management, automation, and continuous service improvement while partnering closely with Application Development, Infrastructure, Cybersecurity, Product, and Business teams.
The role combines deep technical expertise with strong leadership to establish reliability engineering practices, reduce operational risk, improve customer experience, and accelerate delivery through automation and engineering excellence.
In this role, you will:
Reliability & Platform Stability
Own the reliability strategy for critical applications and platforms.
Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
Drive platform availability, resiliency, recoverability, and scalability initiatives.
Reduce Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR) through proactive engineering.
Establish resiliency standards, operational readiness reviews, and production certification requirements.
Incident Management & Problem Management
Lead major incident response activities for high-severity customer-impacting events.
Drive root cause analysis (RCA) and corrective action management.
Establish operational governance and escalation processes.
Identify systemic reliability risks and remediation opportunities.
Coordinate cross-functional recovery efforts involving internal teams and third-party vendors.
Observability & Monitoring
Define enterprise observability standards and best practices.
Drive implementation of:
Metrics
Logging
Tracing
Synthetic Monitoring
Business Observability
Develop operational dashboards and executive-level service health reporting.
Ensure end-to-end visibility across customer journeys and critical business processes.
Automation & Engineering Excellence
Lead initiatives to automate operational processes and reduce manual effort.
Drive adoption of Infrastructure as Code (IaC), CI/CD, Auto-Remediation, and AIOps capabilities.
Improve operational efficiency through self-healing platforms and intelligent alerting.
Eliminate repetitive operational tasks through engineering solutions.
Vendor & Third-Party Reliability Management
Establish operational engagement standards for strategic vendors and partners.
Drive vendor accountability for production stability, incident response, and RCA delivery.
Participate in contractual resiliency reviews and SLA governance.
Ensure third-party technology providers meet operational and reliability expectations.
Risk & Compliance
Partner with Risk, Audit, Cybersecurity, and Compliance teams.
Ensure platforms meet regulatory, resiliency, and operational risk requirements.
Maintain production support controls, procedures, and evidence for audits and examinations.
Required Qualifications:
7+ years of Systems Engineering and Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
3+ years of management or leadership experience
5+ years of leadership experience managing engineering or SRE teams
5 years experience supporting large-scale, customer-facing applications and platforms
5 years experience displaying strong understanding of: Incident Management, Problem Management, SRE Principles, DevOps Practices, Cloud Technologies, Production Operations
Desired Qualifications:
Financial services or highly regulated industry experience.
Expertise with: Splunk, Grafana, AppDynamics, Dynatrace, OpenTelemetry, Prometheus, Kubernetes/OpenShift, Public Cloud Platforms (AWS, Azure, GCP)
Experience implementing Business Observability and AIOps solutions.
Knowledge of ITIL, Operational Resiliency, Disaster Recovery, and Capacity Management frameworks.
Locations:
794 Davis St,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s