Jobs and Careers
BL

Sr. Director of Operations & Reliability

Blue Yonder
Scottsdale, United Statesfull_timeVerifiedPosted 27 Aug 2024
💰 $219,124/yr($143,752/yr$219,124/yr)

About the role

Blue Yonder Title:

Sr. Director of Operations & Reliability

Synonymous Job Title:

Sr. Director, ITG
 

Location:

Scottsdale, AZ preferred
 

Position Overview:

The Sr. Director of Operations & Reliability holds overall accountability for the resiliency and stability of Blue Yonder’s (BY) services, working in close partnership with various divisions across the organization. This role involves designing and establishing BY’s resiliency strategy, including the implementation of reliability engineering best practices across all engineering teams. The Sr. Director will lead incident and problem management efforts across data center (DC) and cloud domains, ensuring thorough root cause analysis and remediation. Additionally, this role will drive operational excellence, minimize downtime, and ensure seamless service delivery across the organization through targeted chaos engineering and automation advancements. This role will report into the CIO.

What you will do:

Operational Leadership:

  • Oversee daily operations to ensure systems and processes are reliable, efficient, and aligned with business goals.
  • Develop and implement operational strategies that enhance performance, scalability, and resilience.
  • Collaborate with cross-functional teams to identify opportunities for process improvement and optimization.
  • Resiliency & Stability: Hold overall accountability for the resiliency and stability of BY’s services, ensuring continuous operation across all business units.

Change Management:

  • Lead the Change Management process to ensure all changes are systematically evaluated, approved, and implemented with minimal disruption.
  • Ensure changes are documented, communicated, and executed according to best practices.
  • Monitor and report on the impact of changes, adjusting strategies as necessary to mitigate risks.

Problem Management:

  • Oversee the Problem Management process to identify root causes of operational issues and implement effective solutions.
  • Lead incident management efforts across DC/Cloud domains, ensuring thorough root cause analysis and remediation.
  • Conduct targeted chaos engineering and testing exercises to validate BY’s resiliency capabilities.
  • Develop and maintain a knowledge base of known issues and solutions to improve response times and reduce downtime.

Reliability Engineering & Strategy:

  • Design and establish BY’s resiliency strategy, incorporating reliability engineering best practices for all BY engineering teams.
  • Utilize data analysis and AI/ML methods to develop engineered solutions that prevent incidents and events within the BY and employee environments.
  • Drive the advancement of automation to reduce manual intervention and error rates.

Risk Management:

  • Identify potential risks associated with operational changes and develop mitigation strategies.
  • Ensure compliance with industry standards, regulatory requirements, and internal policies.
  • Conduct regular risk assessments and implement controls to safeguard operational integrity.

Innovation & Business Strategy:

  • Influence business strategy and outcomes from a technology perspective, ensuring that operational strategies align with business objectives.
  • Partner with various teams across BY to drive innovation through AI/ML and integrate new technologies and processes.

Team Leadership:

  • Lead, mentor, and develop a high-performing team of operations and reliability professionals.
  • Foster a culture of continuous improvement, collaboration, and accountability.
  • Manage team performance through regular feedback, goal setting, and professional development opportunities.

Stakeholder Engagement:

  • Serve as the primary point of contact for operations and reliability issues, working closely with senior leadership and other stakeholders.
  • Communicate operational performance, challenges, and opportunities to executive leadership.
  • Build strong relationships with vendors, partners, and other external stakeholders to ensure service reliability and performance.

What we are looking for:

  • Bachelor’s degree in Operations Management, Engineering, Computer Science, or a related field preferred.
  • Minimum of 10 years of experience in operations, reliability, or related roles, with at least 5 years in a leadership position.
  • Proven experience in Change and Problem Management within a complex, high-availability environment.
  • Strong understanding of reliability engineering principles and practices.
  • Experience in leading incident management,

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Blue Yonder

View company profile →