Jobs and Careers
MA
Senior Director, Head of Platform Reliability and Resilience, US and Bermuda
MarkelUnited Statesfull_timeVerifiedPosted 9 Mar 2026
💰 $259,000/yr($188,000/yr – $259,000/yr)
About the role
What part will you play? If you’re looking for a place where you can make a meaningful difference, you’ve found it.
The work we do at Markel gives people the confidence to move forward and seize opportunities, and you’ll find your fit amongst our global community of optimists and problem-solvers. We’re always pushing each other to go further because we believe that when we realize our potential, we can help others reach theirs. Join us and play your part in something special! About the RoleWe are seeking a senior technology leader to build and lead our Site Reliability Engineering (SRE), Disaster Recovery (DR), and FinOps capabilities across Cloud and Data Center environments.
This role owns the reliability, availability, scalability, resilience, and cost efficiency of all US and Bermuda infrastructure and applications. The successful candidate will establish modern SRE practices, ensure disaster recovery readiness, and drive financial discipline across the platform—partnering closely with Infrastructure, Application, Architecture, Security, and Finance teams. This is a strategic, engineering-led leadership role, not a traditional operations position.
Key Responsibilities
Site Reliability Engineering (SRE)
- Establish and lead an enterprise SRE function across cloud and data center platforms
- Define and operationalize SLOs, SLIs, and error budgets for critical services
- Partner with engineering and application teams to embed reliability into system design and delivery
- Continuously improving service availability, performance, and operational predictability
Observability, Monitoring & Metrics
- Own the enterprise observability strategy (monitoring, logging, tracing, alerting)
- Standardize tools and practices across infrastructure and applications
- Ensure alerts are actionable and aligned to customer and business impact
- Provide executive-level dashboards on platform health, reliability, and risk
Disaster Recovery & Resilience
- Define and own enterprise disaster recovery (DR) strategy, BCP, and execution
- Establish RTO/RPO standards aligned to business criticality
- Ensure DR plans are architected, automated, tested, and audit-ready
- Lead regular DR testing, failover exercises, and resilience reviews
- Partner with Architecture and Security teams to balance resilience, risk, and cost
Reliability Security & Resilience Engineering
- Design and implement security controls that are highly available, scalable, and fault tolerant.
- Collaborate with Security team to Identify and remediate security-related single points of failure (e.g., IAM, secrets management, certificate lifecycles).
- Work with Security team to ensure critical security services (IAM, PKI, WAF, secrets, logging) meet defined SLOs and recovery objectives.
- Partner with Enterprise Architecture to embed Zero Trust and defense-in-depth into resilient system designs.
Reliability Analytics & Data-Driven SRE
- Design and maintain reliability telemetry pipelines across metrics, logs, traces, and events.
- Build and maintain SLO, error budget, and reliability health dashboards aligned to business services.
- Develop and apply predictive analytics and AIOps techniques (anomaly detection, capacity forecasting, alert noise reduction).
Engineering & Platform Reliability
- Own the reliability architecture of cloud platforms across Azure and hybrid environments, ensuring availability, scalability, security, and cost efficiency are designed‑in by default.
- Partner with Cloud Engineering, Architecture, and Application teams to define standard platform patterns for high availability, resiliency, fault tolerance, and multi‑region design.
- Establish and enforce cloud reliability standards, including:
- Resilience patterns (active/active, active/passive, graceful degradation)
- Capacity planning and scalability strategies
- Platform guardrails for reliability, security, and cost controls
- Drive Infrastructure as Code (IaC) and platform automation to reduce configuration drift, manual intervention, and operational risk.
Chaos Engineering & Resilience Testing
- Establish and lead an enterprise Chaos Engineering program to proactively test system resilience and validate failure assumptions across cloud and application platforms.
- Define chaos engineering strategy, principles, and guardrails aligned with business criticality, SLOs, and risk tolerance.
- Partner with SRE, Cloud Engineering, and Application teams to:
- Design and execute controlled failure experiments (infrastructure, network, depende
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s