Jobs and Careers
RI

Senior Staff Software Engineer - Site Reliability Engineering

Ridgeline
San Ramon, United Statesfull_timeVerifiedPosted 9 Jul 2025
💰 $250,000/yr($200,000/yr$250,000/yr)

About the role

Senior Staff Software Engineer - Site Reliability Engineering

Location: Reno, NV; San Ramon, CA

Are you passionate about building systems that make reliability a competitive advantage? Do you enjoy combining hands-on engineering with cross-team influence to improve how organizations operate at scale? Do you thrive in fast-paced environments where you can tackle ambiguity, reduce toil, and experiment with AI-powered tooling? If so, we invite you to join our innovative Site Reliability Engineering team at Ridgeline.

 

As a Site Reliability Engineer at Ridgeline, you’ll be part of a hands-on, strategic team responsible for scaling reliability across our cloud-native platform. You’ll design and improve systems like Health Manager, Incident Command, and observability infrastructure—while also driving forward FinOps tooling and AI-assisted automation that reduce operational burden and surface critical insights. This role is central to Ridgeline’s mission of delivering high-performance, zero-downtime services with speed, clarity, and confidence—and your work will directly empower product, infrastructure, and customer-facing teams to move faster without sacrificing reliability.

 

You must be work authorized in the United States without the need for employer sponsorship.

What will you do?

  • Build and evolve systems like Health Manager, Incident Command, and observability platforms that support zero-downtime deployments and operational readiness
  • Partner with development and infrastructure teams to embed reliability into services and processes
  • Participate in the SRE on-call rotation and lead incident response as needed
  • Design metrics, tooling, and workflows that enable zero-downtime deployments, fast detection, and proactive issue resolution
  • Develop and maintain FinOps tooling to drive cost visibility, usage transparency, and financially-informed engineering decisions
  • Lead incident triage and retrospectives with a blameless, data-driven approach
  • Define observability signals that make system health visible, actionable, and reliable
  • Write production-quality code and ship real improvements—measured by impact, not just effort
  • Drive initiatives that reduce risk, increase visibility, or improve operational resilience across services
  • Foster an outcomes-focused team culture through honest communication, clarity, and accountability
  • Think creatively, own problems, seek solutions, and communicate clearly along the way
  • Contribute to a collaborative environment rooted in learning, teaching, and transparency

 

Desired Skills and Experience

  • 10+ years in software engineering position or similar function, with experience operating large-scale, mission-critical systems
  • Proficiency in one or more of: Kotlin, Java, JavaScript, Python
  • Experience with observability platforms (e.g., Datadog, Prometheus) and monitoring best practices
  • Strong familiarity with infrastructure-as-code tools (e.g., Terraform, CDKTF) and CI/CD systems
  • Experience leading or participating in incident response and service ownership
  • Experience deploying, monitoring, and maintaining multi-tenant architectures
  • Ability to work effectively across teams and communicate technical concepts with clarity
  • Strong written and verbal communication skills, especially in facilitating incident response and working sessions with service teams
  • Comfortable navigating ambiguity and working toward measurable outcomes
  • Proven ability to balance individual contribution with cross-functional impact
  • Experience or interest in FinOps, cost-aware system design, or cloud usage optimization is a plus
  • Familiarity with AI-assisted tooling or workflows is a plus, but not required
  • Willingness to learn about cutting-edge technologies while cultivating expertise in a business domain/problem space.
  • An aptitude for problem solving
  • Ability to communicate effectively
  • Serious interest in having fun at work

 

Who You Are

  • A systems thinker who brings clarity and direction to complex, ambiguous environments
  • A strong communicator who can model transparency, collaboration, and constructive disagreement
  • An engineer who delivers—not just ideas, but real improvements that teams rely on
  • Passionate about outcomes, not just effort—you prioritize what matters and follow through
  • Committed to enabling others by reducing friction, building shared tooling, and simplifying operations
  • Comfortable offering candid feedback and engaging in disagreement with respect and clarity—then committing fully once a decision is made, aligning with the team to drive results

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Ridgeline

View company profile →