Jobs and Careers
OV

Staff Site Reliability Engineer

OVO Group
Any of our officesfull_timeVerifiedPosted 30 Jan 2026
💰 £84,569/yr(£64,070/yr£84,569/yr)

About the role

Role OVO-View

Location: Hub Based - Hybrid

Salary banding:  £64,070 - £84,569

Experience: Mid-level/Expert

Working pattern: Full-Time

Reporting to: Principal Cloud Platform Engineer

Sponsorship: Unfortunately we are unable to offer sponsorship for this role.

This role in 3 words: Automation, Resilience, Observability

Top 3 qualities for this role: Analytical, Proactive, Collaborative

 

Where you’ll work:

Depending on the needs of your business area, we expect hub based people to be in the office at least once a week, and to go to OVO Connection events in-person. 

You’ll be assigned to the closest one of our three hub offices, Bristol, Glasgow, or London; unless your role requires field-based work. Each hub has accessible spaces to park your laptop, is designed to inspire people, help them connect and bring big ideas to life.

 

Everyone belongs at OVO:

At OVO, we are on a mission to solve one of humanity's biggest challenges, the climate crisis. And we know it takes all of us to change the world. That's why we need diverse people from all abilities, gender identities, ethnicities, ages, sexual orientations, life experiences and backgrounds to join us.

 

Teamworking for the planet:

Everything we do here spins around Plan Zero. So, naturally, the team you’ll be joining plays a gigantic role in making that happen. Here’s how:

Site Reliability Engineering is at the heart of OVO's customer-focused technology transformation, building and maintaining scalable, efficient, and reliable platforms for OVO's applications and services. The goal of Site Reliability Engineering is to enhance the reliability, performance, and cost-efficiency of OVO's systems, enabling teams to confidently deliver robust services in GCP. This focus on smart and efficient usage of cloud services also contributes to reducing CO2 usage, which is at the heart of OVO's Plan Zero.

 

This role in a nutshell:

As a Site Reliability Engineer  at OVO, you’ll help ensure our systems are reliable, scalable, and efficient. You’ll focus on maintaining high service availability, improving performance, and optimising how we monitor and respond to incidents. Your expertise in reliability engineering will support continuous improvement, proactively resolve issues before they impact users, and strengthen the overall resilience of our infrastructure.

 

Your key outcomes will be:

  • Developing, Refining, and Automating Monitoring Systems: Design, manage and enhance monitoring, alerting and observability systems - such as Datadog, Prometheus and Grafana - ensuring they deliver meaningful insights and effective alerting. You'll also automate repetitive monitoring tasks to improve efficiency.

  • Managing SLOs/SLIs and Improving Incident Response: Define and track SLOs and SLIs for key services, contributing to better reliability insights. You'll also help refine incident response processes, support on-call operations, and improve tooling and communication during incidents.

  • Incident Management and Post-Mortem Analysis: Play a key role in resolving complex production incidents, leading or supporting technical response efforts. Following incidents, you’ll conduct blameless post-mortems to uncover root causes and drive lasting improvements.

  • Cost Optimisation Implementation: Assess infrastructure usage and apply approved strategies to optimise cloud costs - balancing resource efficiency with performance and reliability.

  • Capacity Planning, Performance Tuning & Resilience: Using monitoring and load testing data, you’ll support capacity planning, recommend performance improvements and help implement resilience best practices across systems.

  • Collaboration and Knowledge Sharing: Work closely with engineering, QA, security and product teams to embed reliability practices, document key processes and mentor peers to support collective learning and growth.

  • Design Review Input:Take part in design reviews, offering guidance on how to improve reliability, scalability and day-to-day operability within system architecture.
  • Community of Practice: Actively contribute to your Community of Practice - leading discussions, sharing experiences, mentoring oth

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

OVO Group

View company profile →