Senior Manager, Site Reliability Engineering
Western Governors UniversityAbout the role
If you’re passionate about building a better future for individuals, communities, and our country—and you’re committed to working hard to play your part in building that future—consider WGU as the next step in your career.
Driven by a mission to expand access to higher education through online, competency-based degree programs, WGU is also committed to being a great place to work for a diverse workforce of student-focused professionals. The university has pioneered a new way to learn in the 21st century, one that has received praise from academic, industry, government, and media leaders. Whatever your role, working for WGU gives you a part to play in helping students graduate, creating a better tomorrow for themselves and their families.
The salary range for this position takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs.
At WGU, it is not typical for an individual to be hired at or near the top of the range for their position, and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is:
Job Description
The Senior Manager of Site Reliability Engineering (SRE) leads the function responsible for ensuring that critical systems and services are reliable, scalable, and resilient. The role combines technical leadership with organizational management, directing SRE teams in designing, implementing, and operating infrastructure that supports business needs. This position defines service reliability standards, drives incident response practices, oversees automation initiatives, and partners with other engineering and product teams to balance reliability with delivery velocity. This position’s main objective is to improve reliability, performance, and operational efficiency to ensure our students and faculty are delighted with the fully online educational experience.
Primary Responsibilities
- Leads and mentors SRE teams, creating an environment that encourages ownership, collaboration, and continuous improvement.
- Establishes the SRE vision, goals, and operational strategies in alignment with organizational objectives.
- Defines reliability roadmaps and communicate priorities to engineering and executive stakeholders.
- Develops, drives, and supports Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) across systems.
- Directs incident management processes, including response coordination, root cause analysis, and follow-up actions.
- Implements practices that reduce downtime and ensure systems meet availability, scalability, and performance expectations.
- Drives adoption of infrastructure as code, CI/CD pipelines, and automated testing to improve operational efficiency.
- Oversees monitoring, alerting, and observability systems that provide insight into service health.
- Evaluates and implements emerging tools that enhance service reliability and reduce manual toil.
- Collects and evaluates system and application data to improve the performance and reliability of the environment proactively.
- Partners with software engineering, security, and product teams to integrate reliability into all development lifecycle phases.
- Provides senior leadership and other stakeholders with transparent reporting on reliability trends, risks, and improvement initiatives.
- Fosters a culture of blameless postmortems and shared accountability for uptime and performance.
- Promotes best practices for resilience, scalability, and disaster recovery.
- Regularly assesses and improves reliability processes and team workflows.
- Stays informed of evolving technologies and practices in SRE, DevOps, AI, Machine Learning, and cloud infrastructure.
- Performs other related duties as assigned.
This job description includes a general representation of job requirements rather than a comprehensive inventory of all required responsibilities or work activities. The contents of this document or related job requirements may change at any time with or without notice.
Qualifications
Knowledge, Skills, and Abilities
- Strong understanding of distributed systems, cloud-native architectures, and infrastructure design.
- Deep familiarity with cloud service providers (AWS, GCP, Azure) and their reliability and security best practices.
- Knowledge of software development lifecycles, DevOps principles, and SRE practices such as SLOs, SLIs, and error budgets.
- Un
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s