Jobs and Careers
GE

Principal Site Reliability Engineer (Intelligent Automation)

Genentech
South San Francisco, United Statesfull_timeVerifiedPosted 13 Apr 2026
💰 $302,000/yr($162,600/yr$302,000/yr)

About the role

A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.

Advances in AI, data and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organizations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The Computational Sciences Center of Excellence (CS CoE) is a strategic, unified group whose goal is to harness the transformative  power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and life-changing  medicines for patients worldwide.

Within the CS CoE organisation, the Data and Digital Catalyst (DDC) organization leads the modernization of our computational and data ecosystems by integrating digital technologies across Research and Early Development to empower stakeholders, advance data-driven science and accelerate decision-making.

The Solutions team within the DDC Organization develops modernized and interconnected computational and data ecosystems.  As a Site Reliability Engineer in the Solutions Engineering capability, you will work closely with our engineering colleagues to  play a pivotal role in designing, implementing, and maintaining scalable, resilient, and supportable cloud-based platform solutions. 

The focus will be on enabling research Application, Machine Learning (ML) workloads and HPC environments through automation, efficient resource management, and Infrastructure as Code (IaC) using tooling.  As a member of the DDC team you will help mature the scalable platforms that help unlock the potential of our diverse scientific data, accelerating the discovery and development of life-changing treatments for patients. 


The Opportunity: 

Infrastructure as Code (IaC) Design and Implementation

  • Architect and implement IaC solutions using tools like Terraform, Spacelift, or CloudFormation to provision and manage cloud infrastructure for ML and HPC workloads. Automate the deployment of scalable ML pipelines, HPC clusters, and supporting services across global regions.

Global Availability and Resiliency

  • Architect resilient and highly available solutions for ML and HPC workloads using cloud-native practices such as auto-scaling, load balancing, and failover mechanisms. Implement disaster recovery (DR) and business continuity plans for critical systems to ensure global operational integrity. Conduct chaos engineering experiments to validate system reliability and identify potential weaknesses.

Automation and Observability

  • Develop automation scripts and workflows to streamline infrastructure management, deployment, and scaling for ML and HPC use cases. Implement robust monitoring, logging, and alerting frameworks using tools like Prometheus, Grafana, Datadog, or ELK Stack to provide deep insights into system health and performance. Knowledge of AIOps incident management, processes and tooling. 

Collaboration and Leadership

  • Provide technical leadership to a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Partner with cross-functional teams to align infrastructure solutions with business objectives and ML/HPC workload requirements. Mentor and train junior engineers in IaC practices, ML, and HPC infrastructure design.

Cost Optimization and Governance

  • Monitor and optimize cloud infrastructure usage and costs for ML and HPC workloads. Ensure compliance with organizational security, governance, and regulatory policies in all IaC and cloud implementations.

Who You Are: 

  • Bachelor’s or Master’s degree in Computer Science or similar technical field, or equivalent experience and 7+ years of experience in software engineering  Site Reliability Engineering (SRE).

  • Proven expertise in supporting and deploying IaC solutions in cloud environments (AWS, Azure, or GCP) for ML and HPC workloads.

  • Background in MLOps pipelines, including model versioning, CI/CD for ML, and feature store integration including experience with managed ML service

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Genentech

View company profile →