Site Reliability Engineer
ONEAbout the role
About One
One’s mission is simple - to help customers achieve financial progress. We’re doing this by creating simple solutions to help our customers save, spend, borrow, and grow their money – all in one place.
The U.S. consumer today deserves better. Millions of Americans today can’t access credit, build savings or wealth, and are left to manage their financial lives through multiple disconnected apps. Almost a quarter of U.S. adults are unbanked or underbanked and roughly 80% of fintech users rely on multiple accounts to manage their finances.
What makes us unique? We are backed by a preeminent fintech investor (Ribbit) and the world’s largest retailer (Walmart), maintain the speed and independence of a startup, and employ a strong (and growing) collection of world-class talent.
There’s never been a better moment to build a business that helps people achieve financial progress. Come build with us!
The role
As a Site Reliability Engineer (SRE) at One, your mandate is to ensure the availability and reliability of our most critical services, and ensure that they meet the requirements of our customers. Our SRE team at One is growing, so you’ll be a crucial early member to help establish the team, processes, and best practices. Success in this role looks like collaborating with other teams to build and run sustainable production systems that can evolve and adapt to the changes in our fast-paced environment.
This role is responsible for:
Working proactively with engineering teams to help them set SLOs and implement best practices for logging and telemetry collection.
Design, implement and maintain the tools and systems that support service reliability, monitoring, and alerting.
Participating in a 12x7 on-call rotation supporting the health of our services.
Driving the incident management process and support a blameless post-mortem culture
Participating in application design consulting and capacity planning.
Defining and formalizing SRE practices and help guide the overall reliability engineering direction.
Providing mentorship both formally and informally to engineers at One.
Continuously optimizing systems and workflows by improving architecture, infrastructure, automation, CI/CD, and observability.
Combining software and systems knowledge to engineer high-volume distributed systems in a reliable, scalable, and fault-tolerant manner.
You bring
10+ years of relevant industry experience with a focus on distributed cloud native systems design, observability, operation, maintenance, and troubleshooting.
5+ years operational experience with an observability platform like Datadog, Splunk, Prometheus/Grafana, or AppDynamics.
Fluency in one or more programming languages (e.g. Python, Typescript, Go).
A strong conviction in software development best practices, including version control, automated testing, and continuous integration and delivery.
You're self-motivated, inquisitive, and always looking to learn new technologies.
You’re a great teammate who communicates clearly and transparently.
The Triple H Factor: Humble, Hungry and Honest.
An act-like-an-owner mentality. We have a bias toward taking action.
Pay Transparency
The estimated annual base salary for this position ranges from $170,000 to $210,000. Pay is generally based upon the level, complexity, responsibility, and job duties / requirements of the specific position. We then source candidates with the requisite skills, expertise, education, training, and experience. If you are selected for an interview, please feel welcome to speak to a Talent Partner about our compensation philosophy and other available benefits.
What it’s like working @ One
Competitive cash
Benefits effective on day one
Early access to a high potential, high growth fintech
Generous stock option packages in an early-stage startup
Remote friendly (anywhere in the US) and office friendly - you pick the schedule
Flexible time off programs - vacation, sick, paid parental leave, and paid caregiver leave
401(k) plan with match
We use Covey as part of our hiring process for jobs in NYC and certain features may qualify it as an AEDT. As part of the evaluation process we provide Covey with job requirements and candidate submitted applications. We began using Covey Scout for Inbound on May 31, 2024.
Please see the independent bias audit report covering our us
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s