Jobs and Careers
WO

Senior Resiliency Lead - DevOps

Workday
United Statesfull_timeVerifiedPosted 13 Feb 2025
💰 $206,500/yr($116,300/yr$206,500/yr)

About the role

Your work days are brighter here.

At Workday, it all began with a conversation over breakfast. When our founders met at a sunny California diner, they came up with an idea to revolutionize the enterprise software market. And when we began to rise, one thing that really set us apart was our culture. A culture which was driven by our value of putting our people first. And ever since, the happiness, development, and contribution of every Workmate is central to who we are. Our Workmates believe a healthy employee-centric, collaborative culture is the essential mix of ingredients for success in business. That’s why we look after our people, communities and the planet while still being profitable. Feel encouraged to shine, however that manifests: you don’t need to hide who you are. You can feel the energy and the passion, it's what makes us unique. Inspired to make a brighter work day for all and transform with us to the next stage of our growth journey? Bring your brightest version of you and have a brighter work day here.

At Workday, we value our candidates’ privacy and data security.  Workday will never ask candidates to apply to jobs through websites that are not Workday Careers. 

  

Please be aware of sites that may ask for you to input your data in connection with a job posting that appears to be from Workday but is not.

  

In addition, Workday will never ask candidates to pay a recruiting fee, or pay for consulting or coaching services, in order to apply for a job at Workday.

About the Team

As a Senior DevOps Engineer you will be joining one of Workday’s most exciting product and technology teams, Core Software. With nearly 750 employees globally, Core Software is responsible for evolving the core technology and runtime components of the Workday Platform and empowering developers to innovate and build on our products. This DevOps Engineer will be reporting to the VP of Core Services, will support our growth and continued success by ensuring our systems are resilient and capable of withstanding various challenges.

We are looking for a Senior Software DevOps Engineer who has proven experience to lead efforts to embed resilience and fault tolerance within our software architecture. Working closely with engineering, operations, and DevOps teams, this role will be responsible for developing and implementing strategies that improve system reliability, reduce Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR), and ensure our services can withstand disruptions.

About the Role

  • Design and deploy fault-tolerant and resilient architecture patterns to improve system availability and performance

  • Facilitate FMEA workshops to identify and prioritize potential failure points within critical applications and services
  • Establish and track metrics such as MTTD, MTTR, and Service Level Objectives (SLOs) to measure and improve system resilience
  • Work with QA and Perf teams to implement chaos engineering, stress testing, and other resiliency testing methodologies
  • Collaborate with Site Reliability Engineering (SRE) and operations teams to define incident management playbooks and recovery procedures for high-severity incidents
  • Educate development teams on best practices for building resilient applications, providing guidance on design principles, patterns, and tools
  • Lead post-incident reviews to identify root causes and implement changes to prevent recurrence, creating a culture of learning and continuous improvement
  • Provide program management support for short- and long-term work related to resiliency, quality, and security
  • Drive priorities and address backlogs for ongoing technical work that will advance reactions to incidents, security improvements, version updating, observability enhancements, and long-term system stability and scalability
  • Support Engineering and Product leaders within the Core Services pillar in establishing plans, managing dependencies and risks, and providing organizational visibility and team accountability to ensure the initiative moves forward at the pace needed
  • Implement techniques and best practices in processes and methodology so teams stay aligned and on track to achieve their stated goals.
  • Indepth OMS and workday stack knowledge is a big bonus

About You

Basic Qualifications

  • 5+ years in software engineering, architecture, or DevOps with a focus on reliability, resilience, and high availability
  • Proficiency in cloud platforms (AWS, GCP), distributed systems, containerization (Kubernetes, Docker), and infrastructure as code
  • Strong knowledge of FMEA, Chaos Engineering, Disaster Recovery, and redundancy strategie

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Workday

View company profile →