Jobs and Careers
GA

Senior Director - Site Reliability Engineering

Gartner
Spainfull_timeVerifiedPosted 7 May 2025

About the role

About Gartner IT: 

Join a world-class team of skilled engineers who build creative digital solutions to support our colleagues and clients.  We make a broad organizational impact by delivering cutting-edge technology solutions that power Gartner.  Gartner IT values its culture of nonstop innovation, an outcome-driven approach to success, and the notion that great ideas can come from anyone on the team.  

About the role: 

Do you wear multiple hats in the world of reliability? Are you a rockstar architect with a passion for building resilient, performant, and scalable systems? If you're a technical leader who thrives at the intersection of software, platform, and on-premise Infrastructure, then we want to meet you!


We are seeking a seasoned Sr. Director of Site Reliability Engineering (SRE) who is responsible for developing and implementing a comprehensive strategy for site reliability, encompassing scalability, performance, and reliability improvements. The role will align SRE objectives with overall conference business goals and technology roadmaps. It will foster the spirit of continuous improvement to the SRE and position it to meet Global Conferences MCPs.
 

The person in this role is responsible for overseeing SRE team operations, ensuring the reliability and availability of all the technical services designed to deliver the world-class experience to our clients. This role will work effectively with Service Management to enforce best practices for system reliability, monitoring, capacity planning, incident response, problem management, disaster recovery, change management, and workflow automation. They will also partner with Global SRE Practice Leader to implement standard tools and technologies necessary to generate a complete view of SRE metrics and improvement areas, including (but limited to) monitoring, logging, notification, dashboarding, and AIOps.

What you’ll do: 

  • Lead three regional cross-functional SREs teams, providing leadership, mentorship, and clear talent strategy to improve the overall effectiveness and performance of SRE team.

  • Design and implement a holistic reliability strategy that aligns with Gartner’s Conferences Business Objectives and Technology Roadmap; encompassing resiliency engineering, performance optimization, and platform stability.

  • Partner with executive leadership to communicate SRE initiatives, advocate for optimal ROI, and drive organizational change.

  • Partner with Business and Technology leaders to develop SLOs and SLAs aligned with business goals.

  • Champion a culture of automation, tooling, and continuous improvement within the reliability team to reduce the toil.

  • Architect and build highly scalable, secure, and resilient digital products ( .NET/NodeJs/Java-based) on-premise and cloud environments such as AWS and Azure with multi-region DR strategy.

  • Implement best practices for observability, leveraging tools and techniques to gain deep insights into system health and performance.

  • Establish 24x7 Incident Management process in partnership with Gartner’s NOC; Oversee incident response, root cause analysis, and post-mortem procedures for rapid resolution and prevention

  • Responsible for tracking and Improving MTTI, MTTR and MTTF for all the technology services, 

  • Stay up-to-date on the latest trends and technologies in reliability, performance engineering, frameworks (such as .NET/Java/Angular..etc), and cloud computing (AWS and Azure).

  • Manage the reliability budget and resources effectively.

  • Advocate for reliability best practices across the entire organization.

  • Travel to the Premium destination conferences to provide onsite support.

What you’ll need:

  • Formal training or certification on site reliability/software engineering concepts and 12+ years applied experience. In addition, 5+ years of experience leading technologists to manage, anticipate and solve complex technical items within your domain of expertise and more broadly across the organization.

  • Minimum 7+ experience leading complex projects supporting site reliability engineering design, scaling, resilience, and system performance assessments for highly critical and regulated applications

  • Strong Technical expertise in .NET/ Angular/ SQL /NoSQL technologies

  • Extensive experience with Cloud platforms (AWS and Azure), including its core

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Gartner

View company profile →