Jobs and Careers
DA

Staff Site Reliability Engineer

Dave Operating LLC
United States, United StatesRemotefull_timeVerifiedPosted 9 Mar 2026
💰 $330,000/yr($208,000/yr$330,000/yr)

About the role

Dave vs. Goliath. We’re Dave.

Dave is a financial app on a mission to build products that level the financial playing field. It is redefining the financial landscape by leveraging technology to create an affordable, transparent, and user-centric access to liquidity for millions of Americans. As a leading innovator in the U.S. financial services sector, Dave’s digital financial platform offers products designed to meet the credit needs of those underserved by traditional financial institutions. Dave’s offerings include its flagship ExtraCash product, providing members up to $500 in short-term advances within minutes. The company is on track to launch several new product offerings in 2026, including a Buy Now Pay Later (BNPL) option.

Dave is focused on serving Americans who are financially vulnerable or living paycheck to paycheck. Dave is leading the charge in creating a new era of credit products that prioritizes speed, affordability, and accessibility, making it the go-to financial partner for those who need it most.

The Opportunity

This is a senior, deeply hands-on role on a small, high-leverage SRE team (3–4 engineers). You’ll serve as a technical anchor across cloud infrastructure and networking, shaping how reliability, automation, and performance are embedded into every layer of our platform.

You won’t just respond to incidents. You’ll design the systems that prevent them. You’ll partner closely with the Director of DevX & Infrastructure Engineering and cross-functional teams to evolve our GCP platform in ways that support product velocity while protecting long-term durability.

What You’ll Build and Own

  • Lead architecture and automation across our GCP environment, ensuring reliability, scalability, security, and thoughtful cost management.

  • Define and improve SLIs, SLOs, and error budgets using Cloud Monitoring and Datadog — connecting reliability goals to real business outcomes.

  • Shape our multi-region, disaster recovery, and capacity planning strategies so the platform holds up as we grow.

  • Design and optimize cloud networking, including VPC architecture, ingress/egress, Cloud Armor, VPN, and DNS to support internal systems, partner integrations, and member-facing services.

  • Drive infrastructure-as-code and GitOps practices using Terraform, Kubernetes, Helm, and ArgoCD to make deployments predictable and repeatable.

  • Mentor SREs and infrastructure engineers through design reviews, incident retros, and hands-on collaboration — strengthening technical depth across the team.

You’ll also explore practical LLM-driven automation where it meaningfully reduces operational toil and shortens incident resolution time.

The Impact

Reliable systems mean members can access ExtraCash, banking, and credit-building tools when they need them most. Your work directly supports trust, growth, and long-term platform resilience.

What we’re looking for

Experience & Technical Foundation

  • 8+ years in software, infrastructure, or site reliability engineering.

  • 5+ years of hands-on experience operating production systems in GCP (compute, networking, storage, IAM, observability).

  • Deep experience with Kubernetes (GKE), Helm, containerization, Terraform (IaC), and ArgoCD.

  • Strong programming skills in Python, Go, or TypeScript/JavaScript for automation and internal tooling.

  • Experience defining and operating against SLIs, SLOs, and error budgets.

  • Strong knowledge of relational and distributed databases (e.g., MySQL, Cloud SQL, Cloud Spanner, Redis), including performance tuning and HA strategies.

  • Experience leading incident response, root cause analysis, and systemic remediation.

Bonus

  • Experience in fintech or regulated environments

  • Familiarity with CI tooling (GHA, Jenkins, Tekton, CircleCI)

  • Experience in high-growth startups.

What Makes Someone Successful Here

You take responsibility for outcomes, not just deliverables. You think in systems — how networking decisions affect latency, how reliability targets affect member trust, how cost decisions affect long-term sustainability. You balance urgency with durability and make thoughtful trade-offs that hold up over time.

You’re comfortable operating in ambiguity. Not everything is fully defined — and you help shape it. You ask good questions before committing to a path and connect day-to-day execution to broader business goals.

You also collaborate naturally. You partner with product engineers, TPMs, and secu

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Dave Operating LLC

View company profile →