Director, Network Service Reliability Engineering
EquinixAbout the role
Who are we?
Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet.
A place where tech thinkers and future builders turn bold ideas into breakthrough experiences, we welcome your unique perspective.Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.Job Summary
The Director, Network Service Reliability Engineering (NSRE) leads the strategy, execution, and evolution of highly reliable, autonomous network services across Equinix. This role is accountable for improving service reliability, scalability, and operational efficiency through autonomous network operations, SRE principles, and data-driven decision making. This role champions a customer-first reliability culture, ensuring every incident becomes an opportunity to strengthen trust through transparency, learning, and durable engineering improvements.
Responsibilities
Network Reliability & Autonomous Operations
Define and execute the global NSRE strategy aligned to Equinix’s business and customer outcomes
Drive adoption of SRE practices (SLIs, SLOs, error budgets) across network services
Lead autonomous network operations initiatives, including AI Ops, event correlation, auto‑incident creation, and self‑healing remediation
Own availability, performance, and resiliency outcomes for mission‑critical network platforms
Operational Excellence & Incident Leadership
Establish and govern reliability KPIs (availability, MTTD, MTTR, change success rates)
Serve as the executive escalation point for high‑severity, customer‑impacting incidents
Lead root cause analysis (RCA), Correction of Errors (COE), and systemic risk reduction
Drive post‑incident learning and ensure durable corrective actions
Engineering, Platforms & Delivery
Partner with Product and Engineering teams to embed reliability, observability, and operability into network design
Promote infrastructure‑as‑code, intent‑based networking, and policy‑driven operations
Oversee a global portfolio of reliability and automation initiatives, ensuring execution against commitments
Remove organizational and technical barriers to delivery
People Leadership & Capability Transformation
Lead and develop global teams across SRE, network reliability, and service management (~80–90 staff)
Build a culture of ownership, accountability, learning, and customer focus
Evolve traditional NOC/SMC capabilities into modern SRE and autonomous operations models
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s