Jobs and Careers
MA

Sr. Manager, Site Reliability Engineering and Cloud Ops

Markel
United Statesfull_timeVerifiedPosted 11 Aug 2025

About the role

What part will you play? If you’re looking for a place where you can make a meaningful difference, you’ve found it.

The work we do at Markel gives people the confidence to move forward and seize opportunities, and you’ll find your fit amongst our global community of optimists and problem-solvers. We’re always pushing each other to go further because we believe that when we realize our potential, we can help others reach theirs. Join us and play your part in something special! We are seeking a highly experienced and strategic Sr. Manager of Cloud Operations & Site Reliability Engineering (SRE) Eo lead our cloud infrastructure initiatives, primarily in Azure and AWS. This senior leader will be responsible for managing cloud operations, ensuring high availability, and optimizing performance through Site Reliability Engineering (SRE) principles.

The ideal candidate will have a strong background in Azure cloud architecture, cloud adoption frameworks, Operations and engineering, DevSecOps, and CI/CD practices, as well as proven success in cost optimization, monitoring, and cloud governance.

Key Responsibilities: 

  • Lead daily operations of cloud platforms (Azure & AWS), ensuring optimal performance, security, scalability, and availability. 

  • Drive implementation of Infrastructure as Code (IaC) and automation for provisioning, configuration, and scaling. 

  • Apply the Microsoft Cloud Adoption Framework to guide enterprise cloud migration and maturity. 

  • Periodic monitoring of health of applications and enforce best practices of Azure landing zone. 

  • Collaborate with cross-functional teams to define and implement best practices around cloud governance, security, and lifecycle management. 

  • Establish and promote SRE principles including automation, observability, incident response, and continuous improvement. 

  • Implement robust monitoring and alerting solutions using tools such as Datadog, Dynatrace, LogicMonitor, and Azure Monitor. 

  • Perform root cause analysis (RCA) and design proactive solutions for system reliability and performance 

  • Manage cloud budgets and costs, optimize spend across Azure and AWS using tools like cost explorer, reserved instances, autoscaling, and right-sizing. 

  • Work closely with DevOps teams to strengthen CI/CD pipelines, enforce security (DevSecOps), and promote end-to-end automation. 

  • Provide strategic input on tooling and architecture for code deployment, containerization, and infrastructure orchestration. 

  • Communicate complex cloud strategies and technical details effectively to both technical and non-technical stakeholders. 

  • Collaborate with internal clients to tailor cloud solutions that align with business goals and compliance standards. 

  • Mentor, grow, and lead a team of cloud engineers and SRE professionals. 

  • Define KPIs and drive a culture of accountability, innovation, and continuous improvement. 

Required Qualifications: 

  • 5–7 years of experience in cloud engineering, cloud operations, and SRE roles. 

  • Proven leadership experience in enterprise cloud environments, preferably at scale. 

  • Deep technical expertise in Microsoft Azure and understanding of AWS. 

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Markel

View company profile →