Senior DevOps Engineer (EKS/Kubernetes)
Caris Life SciencesAbout the role
At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.
We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.
But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.
Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.
Position Summary
The Sr DevOps Engineer is responsible for designing, implementing, and operating secure, scalable Linux-based infrastructure across on-premises and cloud environments, with a primary focus on architecting and running production workloads on Kubernetes/AWS EKS. This role centers on Kubernetes/AWS EKS cluster design, automation, reliability, and compliance initiatives, supporting modern DevOps practices including Infrastructure as Code, CI/CD, and containerization at scale.
Job Responsibilities
Design, deploy, and maintain Linux infrastructure in on-premises and cloud
environments.Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation.
Manage and optimize AWS environments with a focus on performance, scalability, security, and cost efficiency.
Implement and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch).
Architect, deploy, and operate production Kubernetes/AWS EKS clusters, including node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs.
Define and lead cluster upgrade, security hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at scale, while serving as a senior technical resource for complex production incidents.
Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik).
Implement IAM Roles for pod security standards, and network policies to secure EKS
workloads.Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload
autoscaling (HPA/VPA) to optimize performance and cost.Build and maintain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads.
Manage Docker container builds and registries in support of EKS-based application
deployment.Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability.
Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability.
Ensure systems meet security and compliance requirements, including SOX and SOC 2
initiatives.Execute and maintain Linux patching strategies, addressing security updates and CVEs in a timely manner.
Participate in incident response, root cause analysis, and recovery efforts.
Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards.
Participate in on-call rotations and provide after-hours support as required.
Required Qualifications
Bachelor’s degree in computer science, Information Technology or related field.
8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability
Engineering roles.
5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and
CloudWatch.
5+ years of hands-on experience designing and operating production workloads on Kubernetes/AWS EKS, including cluster upgrades, networking, and autoscaling.
Proficiency in scripting and automation using Python and Bash.
Strong hands-on experience with Infrastructure as Code using Terraform, and with CI/CD pipelines (GitLab CI/CD), including running CI/CD workloads on Ku
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s