Senior Manager, Reliability Engineering
ZocdocAbout the role
Our Mission
Healthcare should work for patients, but it doesn’t. In their time of need, they call down outdated insurance directories. Then wait on hold. Then wait weeks for the privilege of a visit. Then wait in a room solely designed for waiting. Then wait for a surprise bill. In any other consumer industry, the companies delivering such a poor customer experience would not survive. But in healthcare, patients lack market power. Which means they are expected to accept the unacceptable.
Zocdoc’s mission is to give power to the patient. To do that, we’ve built the leading healthcare marketplace that makes it easy to find and book in-person or virtual care in all 50 states, across +200 specialties and +12k insurance plans. By giving patients the ability to see and choose, we give them power. In doing so, we can make healthcare work like every other consumer sector, where businesses compete for customers, not the other way around. In time, this will drive quality up and prices down.
We’re 15 years old and the leader in our space, but we are still just getting started. If you like solving important, complex problems alongside deeply thoughtful, driven, and collaborative teammates, read on.
Your Impact on our Mission:
Zocdoc is looking for a Senior Manager, Reliability Engineering to help drive the reliability, resiliency, observability, availability, and scalability of our systems and services. You’ll be challenged to drive continuous improvements to uptime and performance for our patients and providers in a constantly evolving environment. You’ll work with a team of skilled engineers across our AWS cloud-based environments, distributed systems, monolith, and microservices to inform, prioritize, and execute on the most important initiatives. We’re looking for someone who loves challenging the status quo and strives to make everything they touch easier, faster, and more robust.
You’ll enjoy this role if you are…
- A technical leader who thrives when building and coaching engineering teams
- Passionate about ensuring complex systems never skip a beat
- Pragmatic and data-driven in your decision making day-to-day
- Motivated to learn new technologies, design patterns, and work in the cloud
- Comfortable managing failures and outages and evangelizing blameless post-mortems
- Excited to work in a highly collaborative environment with diverse individuals
- Autonomous, individually accountable, and comfortable working in a remote environment
- A believer that diverse and inclusive teams and cultures are non-negotiable
Your day to day is…
- Providing technical and career leadership to a multidisciplinary group of Site Reliability, Database Reliability, and Cloud Engineers
- Scoping, planning, and executing team tasks and projects focused on reliability, resiliency, observability, availability, and scalability of ZocDoc systems and services
- Establishing and measuring clear success metrics for individuals, teams, and projects
- Demonstrating operational excellence throughout the incident management lifecycle
- Serving as a strategic partner and resource to product engineering teams
- Creating and maintaining roadmaps that tie team priorities to business initiatives
- Building scalable and effective business processes which support reliability goals
- Serving as the DRI for specific engineering tools, vendors, and services
- Working across the organization on Engineering culture, helping to build career ladders, steer cross-team events, and build an inclusive technology workforce
You’ll be successful in this role if you have…
- 10+ years of progressive experience in Site Reliability, Database Reliability, or adjacent disciplines (DevOps, Platform, Cloud Engineering, Backend Engineering, etc.)
- 5+ years of experience leading, building, growing, and inspiring engineering teams
- Passion for building and maintaining reliable, resilient, observable, available, and scalable systems and services
- Humility, a strong bias toward action, and exceptional accountability
- Ability to lead and manage critical incident response and retrospectives
- Hands-on experience managing critical, customer facing systems and services
- Demonstrated success in scoping, planning, executing, and measuring team success
- Deep familiarity with distributed systems, edge technologies, cloud services, monolithic, and microservice architectures
- Exceptional written and verbal communication skills that support effective collaboration, negotiation, and influence across technical and non-technical stakeholders
Be
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s