Senior Manager, Site Reliability Engineering
Royal Caribbean GroupAbout the role
Journey with us! Combine your career goals and sense of adventure by joining our exciting team of employees. Royal Caribbean Group is pleased to offer a competitive compensation and benefits package, and excellent career development opportunities, each offering unique ways to explore the world.
Do you thrive in fast-paced environments, leading teams to ensure top-tier website performance and reliability? Do you have a passion for e-commerce and the technical expertise to optimize a complex, high-volume platform? If so, then we have the perfect opportunity for you!
The Royal Caribbean Group’s Global E-Commerce Team has an exciting career opportunity for a full time Senior Manager, Site Reliability Engineering reporting to the AVP, E-Commerce Technology.
Position Summary:
We seek a highly skilled and experienced Senior Manager to lead our Site Reliability Engineering (SRE) function and ensure the optimal performance and availability of our Globa e-commerce platform, which drives over a billion dollars in annual sales. You will oversee a 24x7 monitoring and critical problem resolution function, driving continuous improvement in incident identification and resolution times. Additionally, you will champion consistent measurement approaches for site experience performance and component availability, fostering a culture of observability and reliability across product teams. This role demands exceptional leadership, communication, and problem-solving skills, along with a deep understanding of SRE principles and relevant technologies.
Responsibilities:
-
Oversee 24x7 Monitoring and Critical Problem Resolution:
-
Lead and optimize the effectiveness of a primarily outsourced 24x7 monitoring and critical problem resolution function.
-
Implement strategies to reduce incident identification and resolution times, enhancing overall system reliability.
-
Collaborate closely with managed service providers to define and monitor explicit SLAs and implement continuous improvement tactics.
-
Champion Site Experience Performance Measurement:
-
Establish and implement a consistent approach to measuring site experience performance using industry-standard methodologies (e.g., Google Lighthouse).
-
Analyze performance data and identify areas for optimization, collaborating with product and engineering teams to implement improvements.
-
Drive Component Availability and Observability:
-
Develop and maintain a framework for measuring the availability of various website components owned by different product teams.
-
Foster a culture of ownership and accountability for component reliability and observability across product teams.
-
Ensure World-Class Incident Response:
-
Lead and orchestrate incident response efforts, ensuring clear and transparent communication with stakeholders throughout the process.
-
Implement best practices for incident management, including escalation procedures, communication protocols, and post-incident reviews.
-
Lead Root Cause Analysis and Continuous Improvement:
-
Facilitate comprehensive post-incident root cause analysis to identify underlying issues and prevent future occurrences.
-
Drive the implementation of corrective actions and track their completion to ensure continuous improvement in system reliability.
Qualifications:
-
Bachelor's degree in Computer Science, Engineering, or a related field (Master's degree preferred).
-
7+ years of experience in Site Reliability Engineering or a related field, with at least 3 years in a leadership role.
-
Extensive experience managing and optimizing 24x7 monitoring and incident response functions.
-
Deep understanding of SRE principles, methodologies, and best practices.
-
Strong knowledge and experience with the following technologies:
-
Content Management Systems: Adobe AEM
-
Programming Languages: Java, TypeScript, Javascript
-
Cloud Platforms: AWS (including Lambda, CloudWatch)
-
Containerization: Docker, Kubernetes
-
Front-End Development: React, NextJS
-
Monitoring & Observability Tools: AppDynamics, Splunk, Content Square
-
Incident Management: PagerDuty
-
Excellent leadership, communication, and interpersonal skills.
-
Experience defining and leading outcomes-based agreements with service providers
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s