Jobs and Careers
GL

Senior Manager, Site Reliability Engineering (SRE)

Global Healthcare Exchange (GHX)
Office Location or Remote - USA, United StatesRemotefull_timeVerifiedPosted 3 Dec 2025
💰 $191,000/yr($143,000/yr$191,000/yr)

About the role

 

The Senior Manager, Site Reliability Engineering (SRE) will lead the SRE organization to deliver reliable, scalable, and resilient platforms and services. This role will own the strategy, implementation, and continuous improvement of a unified observability platform that provides end-to-end visibility into infrastructure, applications, APM, and databases, enabling proactive issue detection, faster incident resolution, and improved customer experience. 

The Sr. Manager will drive practices around SLIs, SLOs, SLAs, and Error Budgets, embedding reliability into engineering culture. They will oversee incident management, RCA, proactive alerting, predictive analysis, and automation, while ensuring close collaboration with engineering, product, and platform teams. 

 

Key Responsibilities 

Leadership & Team Management 

  • Hire, lead, and mentor a high-performing SRE team across geographies. 
  • Define and execute the SRE vision, roadmap, and strategy in alignment with business and engineering objectives. 
  • Establish a healthy 24x7 on-call model, ensuring coverage while promoting team well-being. 
  • Drive a blameless culture through structured postmortems and RCA follow-up actions. 

Unified Observability & Monitoring 

  • Build and manage a unified observability platform leveraging tools such as New Relic, Datadog, CloudWatch, Prometheus, Grafana, Graylog, and OpenTelemetry. 
  • Deliver holistic monitoring across infrastructure, applications, databases, APIs, and end-user experience. 
  • Implement APM (Application Performance Monitoring) to trace performance across distributed systems. 
  • Establish dashboards, metrics, and proactive alerting to identify anomalies early. 
  • Drive adoption of AIOps and predictive analytics for proactive reliability improvements. 

Reliability Engineering 

  • Define and manage SLIs, SLOs, SLAs, and Error Budgets across services. 
  • Partner with engineering teams to balance velocity with reliability, ensuring adherence to Error Budgets. 
  • Reduce MTTD (Mean Time to Detect) and MTTR (Mean Time to Resolve) through automation, faster detection, and better instrumentation. 
  • Perform capacity planning, scalability reviews, and resiliency testing. 

Incident & Problem Management 

  • Lead major incident response, coordinating communications with executives and stakeholders. 
  • Drive root cause analysis (RCA) and implement long-term fixes. 
  • Partner with ITSM teams to align with incident, problem, and change management processes. 
  • Ensure continuous improvement loops from incidents back into observability, automation, and engineering practices. 

Collaboration & Cross-Functional Work 

  • Collaborate with Engineering, Product, Security, Cloud, and DevOps teams to embed SRE practices. 
  • Provide guidance on instrumentation, reliability design, and operational readiness for new services. 
  • Partner with DBAs and data platform teams to monitor database health, replication, query performance, and failover readiness. 
  • Champion reliability as a shared responsibility across development and operations. 

 

Qua

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Global Healthcare Exchange (GHX)

View company profile →