Director, Service Reliability Engineering
Marriott InternationalAbout the role
As a leader of the Service Reliability Engineering - Experience organization, you lead the team responsible for accelerating and automating the flow of operational activities, ensuring the reliability, performance and scalability of our critical digital platforms. This role will be responsible for driving the SRE strategy for our digital experience team, implementing best practices, and collaborating closely with cross-functional teams to enhance our infrastructure, observability, automation, incident/problem management and disaster recovery processes. You’ll play a key role in modernizing Marriott's technology stack, fostering a culture of reliability and improving service availability across our customer-facing enterprise applications and its supporting infrastructure. As Director of SRE, you will have the opportunity to shape the reliability and scalability of mission-critical platforms, enabling seamless guest experiences across our global portfolio of brands. If you're enthusiastic about contributing to a resilient architecture and thrive in a collaborative environment, join us at the forefront of innovation. Be a key player in shaping the future of SRE at Marriott, working alongside like-minded individuals who share your passion for speed, confidence, and efficiency.
Qualifications:
- Undergraduate degree in computer science, software engineering, or a related field (or equivalent experience)
- 10+ years of experience in SRE, devsecops or IT operations
- At least 5 years’ experience in a previous leadership role within SRE, devsecops or IT Operations
- At least five years of experience in the following technologies. Certifications preferred:
- Presentation Management – HTML, CSS, JS, Backbone, Node JS, Android, iOS
- Application Platforms – NGINX, Java, Akana, Play Framework, Tomcat, Docker, Openshift
- Application Data – PostgreSQL, Couchbase, Cassandra
- Integration Services – Apache Kafka, Apache Spark, Akana
- Analytics Platforms – Hadoop, dashDB, Cognos, Tableau
- Security – Forgerock, OpenID, OAUTH, Ping Identity
- Public Cloud – Azure, Google Cloud, AliCloud, Amazon Web Services
- CI/CD – Harness (particularly in context of usage within SRE)
- Experience with test automation
- Working knowledge and proven track record of implementing disaster indifferent architecture
- Experience with CDN and Akamai tools
- Linux/Unix system administration experience
- Proficient in scripting and programming languages (like Python, Go, Bash, Shell)
- Hands on experience with infrastructure as code (like Terraform), container orchestration (like Kubernetes), and reliability automation
- Working knowledge of networking, databases, distributed systems
- Deep knowledge of monitoring, logging and incident response tools (like Dynatrace, Splunk, OpsGenie, BigPanda, Prometheus, etc.)
- Experience implementing and maintaining CI/CD pipelines for large-scale applications
- Experience creating system architectures for disaster recovery implementation and failover during disasters
- Familiarity with AI/ML-driven observability and predictive maintenance techniques
- Experience in creating system architectures e.g. for Disaster recovery implementation and failover during disasters.
- Exceptional problem solving, communication and stakeholder management skills
Competencies:
- Experience leading, mentoring and developing high performing SRE teams
- Experience managing large, cross functional vendor teams
- Experience defining SLOs/SLIs, error budgets, and KPIs to drive accountability and performance
- Ability to foster a culture of continuous improvement
- Proven record of staying ahead of industry trends/informed of emerging technologies to enhance system reliability and efficiency
- Experience in hospitality is preferred
CORE WORK ACTIVITIES:
- Define and execute Marriott’s SRE vision, aligning with business objectives and technology roadmaps
- Build, mentor and lead a high-performing SRE team, fostering a culture of collaboration and innovation
- Establish reliability, observability and automation goals to improve system uptime, performance and scalability
- Partner with engineering, operations and security teams to drive best practices and continuous improvement
- Implement reliability-focused engineering practices, including SLAs, SLOs/SLIs and error budgets
- Design and maintain resilient, scalable and fault-tolerant architectures across cloud and hybrid environments
- Develop strategies to proactively identify and mitigate risks to system performance and availabilit
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s