Senior Engineer, Site Reliability
Royal Caribbean GroupAbout the role
Journey with us! Combine your career goals and sense of adventure by joining our exciting team of employees. Royal Caribbean Group is pleased to offer a competitive compensation and benefits package, and excellent career development opportunities, each offering unique ways to explore the world.
We are proud to be the vacation-industry leader with global brands — including Royal Caribbean International, Celebrity Cruises and Silversea Cruises — the most innovative fleet and private destinations, and the best people. Together, we are dedicated to turning the vacation of a lifetime into a lifetime of vacations for our guests.
The Royal Caribbean Group’s Site Reliability Team has an exciting career opportunity for a full time Senior Engineer, Site Reliability reporting to the Senior Manager.
This position is onsite and based in Miramar, Florida.
Tis position is also not eligible for work authorization sponsorship.
Position Summary:
We are seeking a highly skilled Senior Site Reliability Engineer to own, operate, and continuously mature our enterprise observability platform across one of the most complex hospitality and maritime technology environments in the world. This role is the engineering backbone of RCG’s observability practice — responsible for ensuring deep, reliable system visibility across 950+ applications serving 100,000+ users across Royal Caribbean International, Celebrity Cruises, and Silversea.
You will operate at the intersection of infrastructure, application performance, network intelligence, and AIOps — driving measurable improvements in mean-time-to-detect (MTTD), mean-time-to-resolve (MTTR), and overall service reliability. This is a platform engineering and standards leadership role, not a tool administration position.
Key Responsibilities:
Platform Ownership & Architecture
- Own and evolve the enterprise observability platform spanning Cisco AppDynamics, Splunk, ThousandEyes, and PagerDuty AIOps across AWS and Azure environments.
- Architect and enforce a unified telemetry strategy — metrics, logs, traces, and events — standardized via OpenTelemetry across all application tiers.
- Design and govern telemetry data pipelines including ingestion, filtering, routing, and retention to optimize signal quality and platform cost at enterprise scale.
- Drive full-stack observability coverage across ship and shore environments, including maritime network paths, contact center platforms, and revenue-critical booking systems.
SLIs, SLOs & Reliability Engineering
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for all critical services across RCG’s three brands.
- Build alerting frameworks that minimize noise, surface actionable signals, and integrate cleanly with PagerDuty AIOps on-call workflows.
- Partner with SRE teams to drive MTTR reduction, post-incident observability improvements, and proactive reliability practices.
- Instrument and publish DORA metrics (Deployment Frequency, Lead Time, Change Failure Rate, MTTR) to support engineering productivity and release confidence.
AIOps & Intelligent Detection
- Drive AI-assisted incident detection, anomaly correlation, and root cause analysis using PagerDuty AIOps and Splunk IT Service Intelligence (ITSI).
- Tune and mature ML-based alert grouping and noise suppression models to reduce alert fatigue and accelerate triage.
- Integrate observability signals with ServiceNow ITSM for automated incident creation, enrichment, and closed-loop resolution workflows.
Kubernetes & Cloud-Native Observability
- Enable and govern Kubernetes observability for EKS and AKS workloads — container health, resource utilization, pod-level tracing, and cluster performance.
- Integrate observability instrumentation into CI/CD pipelines (GitHub Actions) to enable deployment-correlated performance analysis.
- Maintain and extend AWS CloudWatch and Azure Monitor integrations to ensure cloud infrastructure is fully represented in the observability estate.
Standards, Enablement & Technical Leadership
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s