Jobs and Careers
BL

Site Reliability Engineer, Consultant

Blue Shield of California
United Statesfull_timeVerifiedPosted 23 Feb 2026

About the role

Your Role 

We are seeking an Experienced Site Reliability Engineer (SRE) to lead reliability, scalability, and performance initiatives across our production systems. In this role, you will blend software engineering, automation, and systems operations to ensure that our platforms are resilient, efficient, and continuously improving.
You will be part of a cross-functional team responsible for designing, implementing, and maintaining reliable systems that support millions of requests daily. This position requires a deep understanding of distributed systems, cloud infrastructure, automation, and incident response.

Your Work

In this role, you will

  • Reliability & Uptime: Design and maintain systems to achieve high availability (99.9%+), scalability, and resilience.
  • Monitoring & Observability: Build and improve monitoring stacks using tools like Prometheus, Grafana, Datadog, or New Relic.
  • Automation: Reduce manual toil by automating deployments, scaling, and recovery processes using IaC (Terraform or CloudFormation).
  • Incident Management: Lead and respond to production incidents, perform root-cause analysis, and drive postmortems and prevention strategies.
  • Performance Optimization: Identify system bottlenecks and improve performance across compute, network, and database layers.
  • Capacity Planning: Forecast growth, conduct load testing, and ensure services can handle future demand.
  • Security & Compliance: Implement best practices for infrastructure security, secrets management, and compliance requirements.
  • Collaboration: Partner with developers to embed reliability practices into the SDLC, CI/CD pipelines, and application architecture.
  • Chaos Engineering: Design and execute chaos testing experiments to proactively identify weaknesses in distributed systems and improve overall resilience.
  • Deployment Strategies: Implement and manage Blue/Green and Canary deployment methodologies to minimize risk and ensure safe, incremental rollouts of new features and updates.

Your Knowledge and Experience

Education & Experience

  • Requires a Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience); Master’s degree a plus.
  • 7+ years of experience in building, supporting, and improving production systems and infrastructure.

Cloud Platforms

  • Minimum 5 years of hands-on experience with Azure, AWS, or GCP.
  • Demonstrated expertise in virtual machines (VMs), containers, cloud networking, identity and access management (IAM), monitoring, storage, and serverless functions.
  • Comfortable deploying and managing cloud-native services and infrastructure.

Programming & Scripting

  • Proficiency in one or more languages such as Python, Go, Java, Bash, PowerShell, or similar.
  • Ability to write clean, maintainable code for automation and tooling.

Containerization & Orchestration

  • Experience working with Kubernetes, Docker, and tools like Helm or Red Hat OpenShift.
  • Familiarity with managing containerized applications in production environments.

Monitoring & Observability

  • Working knowledge of tools such as Prometheus, Grafana, Datadog, New Relic, ELK Stack, Dynatrace, Splunk, Big Panda, SolarWinds.
  • Ability to set up dashboards, alerts, and metrics to ensure system health and performance.

CI/CD & Configuration Management

  • Experience with CI/CD pipelines using tools like Jenkins, GitHub Actions, GitLab CI, Argo CD, Spinnaker.
  • Familiarity with configuration management tools such as Ansible, Chef, Puppet.

Automation & Emerging Technologies

  • Understanding of Agentic AI systems and automation frameworks for incident response and infrastructure optimization is a plus.
  • Interest in exploring intelligent automation to improve reliability and reduce manual toil.

Testing & Deployment Expertise

  • Experience with chaos engineering tools (e.g., Gremlin, Chaos Monkey) and methodologies.
  • Hands-on knowledge of Blue/Green and Canary deployment strategies in cloud-native environments.

#LI-EB1

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Blue Shield of California

View company profile →