Senior Site Reliability Engineer
CVS HealthAbout the role
Bring your heart to CVS Health. Every one of us at CVS Health shares a single, clear purpose: Bringing our heart to every moment of your health. This purpose guides our commitment to deliver enhanced human-centric health care for a rapidly changing world. Anchored in our brand — with heart at its center — our purpose sends a personal message that how we deliver our services is just as important as what we deliver.
Our Heart At Work Behaviors™ support this purpose. We want everyone who works at CVS Health to feel empowered by the role they play in transforming our culture and accelerating our ability to innovate and deliver solutions to make health care more personal, convenient and affordable.
If you’re eager to make a real impact in the health care industry through your own meaningful contributions, explore a role in technology with CVS Health. Our journey calls for technical innovators and data visionaries: come help us pave the way.
This position can be hybrid to any major CVS hub.
As Senior Site Reliability Engineer, you will:
Identify, curate, implement and adapt critical metrics for not only system health and performance but also for team management and success.
• Improve the SLA, which included maximizing operational efficiencies, strengthening incident management, problem management and knowledge sharing practices.
• Practice sustainable incident response and blameless postmortems.
System Support:
• Understand the technology stack end-to-end and ability to keep up with changing non-functional and functional requirement.
• Build unified monitoring framework, develop efficient automation, deliver solutions to improve the reliability of systems.
• Be an advocate of security best practices, champion and support the importance of security within engineering, partnering with Enterprise security teams and product owners to ensure compliance.
• Provide recommendations for continuous improvement. Provide technical leadership direction, determining and developing approaches to solutions by coordinating multiple resources to solve complex problems.
• Authority in infrastructure topology, resiliency patterns and observability, and SRE practices.
• Subject matter expert for multi-cloud configurations, containerization technologies, with a passion for automation and deep knowledge of DevOps
• Work to simplify and automate deployment processes, run-time operations and provide non-disruptive releases.
Required Qualifications:
5+ years of overall IT experience with software engineering/support or systems engineering background.
3+ years of experience with DevOps and SRE principles, including CI/CD pipelines, incident management, infrastructure as code, and proactive monitoring.
3+ years of experience with source control and continuous integration tools: Git/Stash, BitBucket, Jenkins, Dimensions, Changeman, Jira, Rally, etc.
3+ years of experience with logging platforms and application performance metrics: Splunk, ELK, DataDog, AppDynamics, NewRelics, etc
3+ years of experience in databases such as DB2, Oracle, PostgreSQL, MySQL, Mongodb
3+ years of experience with observability tools such as Prometheus, Grafana, DataDog, etc.
Preferred Qualifications:
• Experience programming with Oracle, Splunk, HTML, Nodejs, Java/J2E
• Experience on Cloud Technologies (AWS, Azure, Google), Microservices, Rancher, Docker, Kubernetes, and web APIs
• Knowledge of Application Cloud Security, Networking including DNS, WAF, DHCP, Firewalls and IP routing
• Experience in partnering with architecture, product, and program management teams to influence product development assisting or improving products.
• Experience in operating and implementing distributed and highly concurrent service-based architecture, including microservices, containerized services, serverless architecture.
• Operational experience in complex distributed and real-time systems, including experience in understanding the SLO/SLAs to understand the non-functional requirements associated with high availability, reliability, and DR goals.
• Knowledge of building and developing observability platforms
• Experience in instrumentation with systems skills in building and operating, monitoring, metering, logging, alerting services of distributed systems at scale
• Experience in operating and implementing distributed and highly concurrent service-based architecture, including microservices, containerized services, serverless architecture.
Education:
Bachelors in Engineering or equivalent experience
Pay Range
The typical pay range for this role is:
$92,700.00 - $222,480.00
This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s