Jobs and Careers
CU

Site Reliability Engineer

Cutover
Remote US, United StatesRemotefull_timeVerifiedPosted 16 Oct 2025
💰 $130,000/yr($120,000/yr$130,000/yr)

About the role

An inclusive work environment is an empowering one. At Cutover, we lead with empathy and enable others to succeed through curiosity, kindness, and self-expression.

Location: Remote, United States

2nd Shift: 2:00pm -11:00pm PST  (10:00 PM - 7:00 AM UTC)

Cutover provides enterprise technology operations teams with an AI-powered SaaS solution that automates and streamlines complex processes with intelligent runbooks. The Cutover solution enables teams to respond to incidents quickly, recover from IT outages, and manage cloud migrations with precision and efficiency. Cutover is used in many of the world's largest financial institutions to support their critical technology operations, including 5 out of the top 6 largest asset managers and 3 out of the top 5 US banks.

We’re looking for a Site Reliability Engineer (SRE) to add to our US team. This role will report to our SRE Lead.

Cutover’s SRE team is responsible for ensuring the reliability and performance levels of our production systems and applications. As a team, we’re committed to constantly improving our engineering culture to maintain a balance between risk and reliability.

 

What tech stack do we use here at Cutover?

The platform is built on a ReactJS frontend with a Ruby on Rails API, and all hosted on the reliable infrastructure of Amazon Web Services (AWS).

Your role will involve close collaboration with our support and engineering teams. Together, we actively engage in maintaining and optimizing the platform's reliability, utilizing cutting-edge tools and occasionally leveraging in-house software and scripts.

If you're passionate about ensuring the dependability and efficiency of complex systems and thrive in an environment where technologies like React, Ruby, AWS, Kubernetes, Terraform, Git, and Ansible are at the forefront, we invite you to join our team. Together, let's elevate the reliability of our Cutover Enterprise platform to new heights.

 

As a Site Reliability Engineer, here's what you'll be up to:

  • Incident Response: Respond to incidents and alerts, triaging urgency and investigating root cause
  • Documentation: Regular contributions to improve our documentation on system design, troubleshooting, best practices, and engineering processes
  • Root Cause Analysis: Contribute to post-mortems and help identify long-term improvements under guidance
  • Collaboration: Support cross-functional teams during investigations and post-incident reviews
  • Observability: Support and enhance observability tools and techniques by identifying metrics, logging, and alerting improvements
  • Automation: Write and execute simple automation scripts (e.g. Python, Ruby, Bash) to improve reliability and toil reduction
  • Development: Work on internal tools, pipelines, and IaC solutions to help improve the speed of software delivery and recovery
  • System Reliability: Work on efforts to enhance the reliability and performance of our application and systems, ensuring optimal uptime and minimal disruptions.
  • Infrastructure Optimization: Work closely with the development and platform engineering teams to optimize the infrastructure on AWS, ensuring scalability and efficiency.

Please note that this role involves a rotating on-call schedule, which will require occasional evening and weekend availability.

 

What we'd like you to bring to the table:

  • A genuine excitement for complex problem solving within our tech stack, applying what you know to our unique problems.
  • Familiarity with at least one scripting language such as Ruby, JavaScript, Python, Bash
  • Experience with containerization (i.e. Docker) or IaC (e.g. Terraform, Helm, CloudFormation)
  • An eagerness to follow modern engineering practices and learn from others
  • Familiarity with observability tools such as DataDog, New Relic, Grafana, Prometheus, ELK, or OpenTelemetry
  • Understanding of core networking concepts (DNS, HTTP/S, Load Balancing, etc.)
  • A collaborative mindset with clear communication skills
  • Willing to ask questions to gain a better understanding of new or complex concepts

Nice to haves…

  • Exposure to major incident response processes
  • AWS Certified Cloud Practitioner or hands-on experience with cloud environments

 

The good stuff…

  • We're excited to offer Share Options as part of our compensation package.
  • 20 days of PTO per year + public holidays, and we want you to take all of them!
  • 3 volunteer days to use for any charitable/voluntary cause you would like.
  • A top-tier pr

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Cutover

View company profile →