Sr. Manager, Reliability Engineering
ZocdocAbout the role
Our Mission
Healthcare should work for patients, but it doesn’t. In their time of need, they call down outdated insurance directories. Then wait on hold. Then wait weeks for the privilege of a visit. Then wait in a room solely designed for waiting. Then wait for a surprise bill. In any other consumer industry, the companies delivering such a poor customer experience would not survive. But in healthcare, patients lack market power. Which means they are expected to accept the unacceptable.
Zocdoc’s mission is to give power to the patient. To do that, we’ve built the leading healthcare marketplace that makes it easy to find and book in-person or virtual care in all 50 states, across +200 specialties and +12k insurance plans. By giving patients the ability to see and choose, we give them power. In doing so, we can make healthcare work like every other consumer sector, where businesses compete for customers, not the other way around. In time, this will drive quality up and prices down.
We’re 17 years old and the leader in our space, but we are still just getting started. If you like solving important, complex problems alongside deeply thoughtful, driven, and collaborative teammates, read on.
Your Impact on our Mission:
At Zocdoc, we’re on a mission to give power to the patient—and that means making sure our systems are always available when patients and providers need them. As the Senior Manager of Reliability Engineering, you’ll lead our SRE, DBRE, and Cloud Engineering teams to deliver highly available, scalable, and observable systems. Your leadership will shape the foundational infrastructure that powers every appointment booked, every provider searched, and every connection made.
This is a pivotal role in the Infrastructure organization. You’ll help reduce incidents, optimize performance, and drive operational excellence across our distributed environments, from legacy monoliths to modern microservices. You’ll collaborate across engineering to enable faster delivery, safer deployments, and a more resilient platform—helping Zocdoc meet the needs of millions of users.
You’ll also help define Zocdoc’s internal reliability platform strategy, building scalable internal services that product teams can adopt easily and confidently. Your team will support our ongoing migration to cloud-native architecture while maintaining thoughtful investment in our monolith, balancing modernization with system stability. And you’ll play a central role in cloud tooling and vendor strategy, helping us make high-ROI decisions on infrastructure investments.
You’ll enjoy this role if you are…
- A technical leader who thrives on building and coaching high-performing, mission-driven engineering teams
- Passionate about reliability, observability, and operational rigor in complex, distributed environments
- Driven by impact and pragmatism, with a bias toward data-informed decisions and continuous improvement
- Comfortable navigating failures and ambiguity, and experienced in leading high-stakes incident response
- Motivated to work across infrastructure domains (SRE, DBRE, Cloud) and evolve platform maturity
- Energized by building reliable platforms as internal products and enabling others to move faster with confidence
- A strong advocate for inclusive, empowering team cultures that foster growth and accountability
- Comfortable working autonomously and collaboratively in a distributed and flexible work environment
- Excited to grow as a leader, working closely with senior infrastructure leadership to expand your influence
Your day to day is…
- Leading and mentoring a multi-disciplinary team of SREs, DBREs, and Cloud Engineers through technical execution and career development
- Defining and driving initiatives that improve reliability, resilience, observability, and performance across our platform
- Partnering with Engineering and Product leaders to ensure systems are safe by default and operationally sustainable
- Driving incident management practices—from fast response to meaningful retrospectives that lead to real change
- Creating and evolving metrics-driven processes to track team, service, and project health
- Owning roadmaps that link infrastructure investments to business value, including cost optimization and tooling strategy
- Managing cloud tooling and vendor decisions to maximize ROI and technical leverage
- Supporting hybrid architectural strategies, improving the reliability of both our monolith and cloud-native services
- Acting as a strategic resource for cloud operations, shared platforms, and system-wide reliability architecture
- Championing technical culture
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s