Lead Director, SRE - Production Support
CVS HealthAbout the role
Bring your heart to CVS Health. Every one of us at CVS Health shares a single, clear purpose: Bringing our heart to every moment of your health. This purpose guides our commitment to deliver enhanced human-centric health care for a rapidly changing world. Anchored in our brand — with heart at its center — our purpose sends a personal message that how we deliver our services is just as important as what we deliver.
Our Heart At Work Behaviors™ support this purpose. We want everyone who works at CVS Health to feel empowered by the role they play in transforming our culture and accelerating our ability to innovate and deliver solutions to make health care more personal, convenient and affordable.
Position Summary
The Lead Director of production support and SRE is a highly visible role that will lead and oversee the production support team and activities for the entire Trade, Underwriting and Finance technology organization. In addition to bringing good SRE practices to the organization, the leader will be responsible for standardizing support processes, ensure incident, problem communication, define SLAs, measure, track and publish production support metrics. The leader should have prior experience inheriting teams with non standard practices and bringing in method, structure and standards and production support framework for the application areas. The leader will build out the SRE function. Hire, mentor, coach SRE Engineers.
Responsibilities
Identify, curate, implement and adapt critical metrics for not only system health and performance but also for team management and success.
Operational experience in complex distributed and real-time systems, including experience in understanding the SLO/SLAs to understand the non-functional requirements associated with high availability, reliability, and DR goals.
Experience in building and developing observability platforms
Experience in instrumentation with systems skills in building and operating, monitoring, metering, logging, alerting services of distributed systems at scale
Experience in operating and implementing distributed and highly concurrent service-based architecture, including microservices, containerized services, serverless architecture.
Improve the SLA, which included maximizing operational efficiencies, strengthening incident management, problem management and knowledge sharing practices.
Practice sustainable incident response and blameless postmortems.
Understand the technology stack end-to-end and ability to keep up with changing non-functional and functional requirement.
Build unified monitoring framework, develop efficient automation, deliver solutions to improve the reliability of systems
Be an advocate of security best practices, champion and support the importance of security within engineering, partnering with Enterprise security teams and product owners to ensure compliance
Provide recommendations for continuous improvement. Provide technical leadership direction, determining and developing approaches to solutions by coordinating multiple resources to solve complex problems
Authority in infrastructure topology, resiliency patterns and observability, and SRE practices.
Subject matter expert for multi-cloud configurations, containerization technologies, with a passion for automation and deep knowledge of DevOps
Work to simplify and automate deployment processes, run-time operations and provide non-disruptive releases.
Technically lead and mentor the team with a focus towards improving the availability, reliability and observability IT products while reducing the burden of toil with tooling, automation, or process change.
Deep external and industry expertise and the ability to introduce and shape automation approaches and practices
Establishes a clear vision aligned with company values; sets specific challenging and achievable objectives and action plans; motivates others to balance customer needs, budgets, and business success.
Develops an organization that attracts, selects, and retains high caliber, diverse talent able to successfully achieve or exceed stated goals; builds a cohesive team that works well together and across other technology segment functions
Required Qualifications
10+ years of overall IT experience with software engineering or systems engineering background.
8+ years of direct management experience leading Site Reliability Engineers
Experience in delivering large scale software with modern reliability and resilience concepts (multi-region, multi-cloud, active/active, canary
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s