Site Reliability Engineer
LifeStance HealthAbout the role
At LifeStance Health, we strive to help individuals, families, and communities with their mental health needs. Everywhere. Every day. It’s a lofty goal; we know. But we make it happen with the best team in mental healthcare.
Thank you for taking the time to explore a career with us. As the fastest growing mental health practice group in the country, now is the perfect time to join our team!
LifeStance Health Values
Belonging: We cultivate a space where everyone can show up as their authentic self.
Empathy: We seek out diverse perspectives and listen to learn without judgment.
Courage: We are all accountable for doing the right thing - even when it's hard - because we know it's worth it.
One Team: We realize our full potential when we work together towards our shared purpose.
ROLE OVERVIEW
At LifeStance Health, we’re building the future of mental healthcare—and we need a Senior Site Reliability Engineer to architect and safeguard the mission-critical infrastructure behind our national digital health platform.
This is not just a support role. You’ll be a principal engineer shaping how our platform scales securely and reliably to serve millions. You’ll define service-level objectives (SLOs), lead reliability reviews, champion incident response, and ensure production readiness is embedded in our engineering DNA.
Your work will power clinical care at scale—delivered with precision, performance, and resilience.
COMPENSATION: $140,000 - 160,000/annually, plus annual bonus potential
CHALLENGES & OPPORTUNITIES
• SLO-Driven Engineering – Define and enforce SLIs, SLOs, and error budgets that guide engineering velocity without compromising reliability.
• Autonomous Infrastructure – Architect systems that heal, scale, and upgrade themselves using Terraform, Kubernetes, and GitOps pipelines.
• Platform Reliability at Scale – Design for 99.99% uptime across distributed systems running in highly regulated environments.
• Proactive Incident Engineering – Build chaos engineering, runbooks, and run incident simulations that drive faster MTTR and a blameless culture.
• Observability as First-Class – Advance full-stack observability using distributed tracing, real-time analytics, and SLO dashboards that drive reliability.
• Security by Design – Lead defense-in-depth infrastructure strategies that meet or exceed HIPAA, SOC 2, and industry-grade threat models.
• Mentorship & Influence – Serve as a force multiplier by mentoring engineers, leading production reviews, and evolving the reliability mindset company-wide.
KEY RESPONSIBILITIES
• Architect scalable, secure infrastructure on AWS using EKS, Lambda, and edge networking strategies.
• Define and own SLOs/SLIs for key services; integrate error budgets into product and deployment planning.
• Drive incident response operations, lead postmortems, and institutionalize RCA learnings.
• Automate everything: provisioning, security controls, deployments, chaos, DR drills—using Terraform, Helm, GitHub Actions.
• Build and maintain observability stack (Datadog, Prometheus, ELK, OpenTelemetry); deliver actionable dashboards and alerts.
• Engineer for cost-aware scale: right-size compute, optimize network paths, and containerize performance-hardened workloads.
• Implement and maintain zero-trust IAM and secrets management frameworks (Vault, AWS Secrets Manager).
• Lead platform reliability reviews and collaborate with engineers, security, and compliance teams to harden architecture.
REQUIREMENTS
• 10+ years in DevOps/SRE/Platform Engineering roles; at least 4+ years architecting for distributed cloud-native systems at scale.
• Expert in AWS core services (EKS, VPC, RDS, Route 53, IAM, Lambda); Terraform-first mindset.
• Proven track record in establishing SLIs/SLOs, building error budgets, and aligning them with business velocity.
• Deep expertise in Kubernetes (EKS), Helm, service meshes (Istio/Linkerd), and microservices orchestration.
• Strong software engineering fundamentals in Python, Go, or similar.
• Hands-on experience with modern observability platforms and real-time monitoring solutions.
• Technical leadership in incident response, risk management, and operational resilience in regulated industries.
• Ability to translate system architecture into platform strategy and influence executive stakeholders.
PREFERRED SKILLS & KNOWLEDGE
• Certifications: AWS DevOps Pro, GCP SRE/Architect, Certified Kubernetes Administrator (CKA).
• Experience with hybrid/multi-cloud systems and edge deployments.
• Experience deploying and securing healthcare platf
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s