Jobs and Careers
RU
Staff Software Engineer - Reliability (US Citizen Only)
Rubrik Job BoardPalo Alto, USAfull_timePosted 15 Jul 2026
About the role
<h2><strong>About Team &amp; About Role</strong></h2> <p>The Site Reliability Engineering (SRE) team at Rubrik ensures the absolute reliability, availability, performance, and security of our enterprise infrastructure services, spanning both global SaaS platforms and government-compliant environments. We operate at the intersection of software development and systems engineering, prioritizing hyperscale platform automation, self-healing architectures, and structural resiliency. As a Staff Site Reliability Engineer, you will serve as a primary technical leader and architect across our broader distributed cloud systems. You will drive long-term technical roadmaps, establish cross-organizational reliability standards, and solve complex distributed systems challenges that safeguard both enterprise and public sector environments.&nbsp;</p> <p>Beyond the core SRE charter, this Staff role also leads the Application-SRE team — a US-based group that partners closely with engineering, Sales, and Support to unblock POCs, drive complex customer escalations to resolution, and convert recurring field signals into engineering and reliability roadmap items. You will be the technical leader and project owner for Application-SRE: setting direction, tracking commitments, and ensuring the team operates as a high-leverage bridge between the field and the broader engineering org.</p> <p>&nbsp;</p> <h2><strong>What You'll Do</strong></h2> <p>As a Staff Site Reliability Engineer, you will possess engineering-wide influence and take ownership of the following critical areas:</p> <ul> <li><strong>Infrastructure Strategy &amp; Architecture:</strong> Formulate and execute the architectural vision for Rubrik's Cloud Platform, optimizing backend infrastructure systems like Kubernetes, MySQL, and cloud-native services for performance, security, and multi-region scale.</li> <li><strong>Hyperscale Automation &amp; Platform Tooling:</strong> Build, scale, and maintain sophisticated custom internal tools, platform controllers, and automation frameworks in Go or Python to systematically eliminate operational toil.</li> <li><strong>AI Infrastructure for SaaS:</strong> Deploy, scale, and operate the AI infrastructure that powers Rubrik's SaaS offerings, owning the reliability, performance, cost, and security controls required to run AI workloads in multi-tenant, compliance-bound environments.</li> <li><strong>AI for SRE &amp; Engineering Productivity</strong>: Drive the adoption of AI-driven solutions across the SRE charter to compress toil and multiply the org - applying agentic and LLM-based approaches to automated triage, incident response, operational analysis, and developer productivity.</li> <li><strong>AI Adoption Guardrails for SaaS R
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s