Site Reliability Engineer Jobs in San Francisco, USA

73 verified site reliability engineer openings in San Francisco

  • Staff Site Reliability Engineer, Tech Lead

    unify · San Francisco Office, USA

    On-site
    17 days agoApply →
  • Senior Site Reliability Engineer

    unify · San Francisco Office, USA

    On-site
    17 days agoApply →
  • Senior Staff Site Reliability Engineer, Tech Lead

    unify · San Francisco Office, USA

    On-site
    17 days agoApply →
  • Senior Site Reliability Engineer

    Airbyte · San Francisco, USA

    On-site
    19 days agoApply →

    Senior Site Reliability Engineer at Airbyte. Apply via Ashby.

  • Site Reliability Engineer (SRE)

    airapps · San Francisco, USA

    On-site
    19 days agoApply →
  • Senior Site Reliability Engineer, Spend

    airwallex · US - San Francisco, USA

    On-site
    19 days agoApply →
  • Senior Site Reliability Engineer

    hyperbolic · San Francisco, CA, USA

    On-site
    20 days agoApply →
  • Manager, Site Reliability Engineering

    okta · San Francisco, California

    On-site
    20 days agoApply →
  • Site Reliability Engineer, Frontier Systems Infrastructure

    openai · San Francisco, USA

    On-site
    21 days agoApply →
  • Site Reliability Engineer

    baseten · San Francisco, USA

    On-site
    21 days agoApply →
  • Senior Site Reliability Engineer

    alembic · San Francisco HQ, USA

    On-site
    21 days agoApply →
  • Senior Network & Site Reliability Engineer

    alembic · San Francisco HQ, USA

    On-site
    21 days agoApply →
  • Senior Site Reliability Engineer

    airbyte · San Francisco, USA

    On-site
    21 days agoApply →
  • Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

    Anyscale · San Francisco, USA

    On-site
    26 days agoApply →

    Senior Site Reliability Engineer, Platform Infrastructure (Foundations) at Anyscale. Apply via Ashby.

  • Site Reliability Engineer (Senior or Staff), Infrastructure Security

    MongoDB · Austin; New York City; San Francisco; Seattle; United States, USA

    On-site
    about 1 month agoApply →

    <p>We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.</p> <p>The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture.</p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation.</p> <p>This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones.</p> <h3>Responsibilities:</h3> <p>Cloud Security Design and Implementation:  </p> <ul> <li>Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM)</li> </ul> <p>Automation and Monitoring:  </p> <ul> <li>Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection</li> </ul> <p>Security Tooling:</p> <ul> <li>Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM </li> </ul> <h3>Qualifications:</h3> <p>Experience:  </p> <ul> <li>6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role</li> </ul> <p>Security Mindset:</p> <ul> <li>A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking layers to cloud provider configurations. Proven experience in leading projects within security-focused areas, such as runtime scanning, security observability, CSPM, and more</li> </ul> <p>Cloud Expertise:  </p> <ul> <li>Strong experience with at least one cloud platform (AWS, Azure, GCP), including expertise in IAM, VPC networking, security groups, and cloud security tools (e.g., GuardDuty, Security Hub, CloudTrail)</li> </ul> <p>Coding/Automation:  </p> <ul>

  • Senior Site Reliability Engineer, Fleet Management

    MongoDB · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, USA

    On-site
    about 1 month agoApply →

    <h3>The Team</h3> <p>Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems.</p> <p>The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model.</p> <p>This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin.</p> <h3>Responsibilities</h3> <ul> <li>Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB</li> <li>Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems</li> <li>Participate in a 24/7 on-call rotation to resolve critical issues</li> <ul> <li>Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice</li> </ul> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 6+ years of experience in software development and operating distributed systems</li> <li>Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests)</li> <li>Have deep experience using and extending containerization technologies, preferably Kubernetes</li> <li>Have a solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS)</li> <li>Possess a customer focused mindset, treating internal developers as your primary users</li> <li>Have strong operational ownership, including a track record of debugging complex production issues and driving them to resolution</li> <li>Prefer automation over manual processes ("allergic to ops work")</li> <li>We are a small team of software engineers with a strong bias toward building software solutions to eliminate toil</li> </ul> <h3>Strong candidates may also have experience w

  • Site Reliability Engineer, Frontier Systems Infrastructure

    Openai · San Francisco, USA

    On-site
    about 1 month agoApply →

    Site Reliability Engineer, Frontier Systems Infrastructure at Openai. Apply via Ashby.

  • Site Reliability Engineer

    Latent · San Francisco, USA

    On-site
    about 1 month agoApply →

    Site Reliability Engineer at Latent. Apply via Ashby.

  • Site Reliability Engineer II, tvScientific

    pinterest · San Francisco, US

    Remote
    about 1 month agoApply →
  • Senior Site Reliability Engineer

    You.com · San Francisco (), USA

    Hybrid
    about 1 month agoApply →

    <div class="content-intro"><p><span style="text-decoration: underline;"><strong>About Us</strong></span></p> <p>At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-time, accurate, and citation-backed information.</p> <p>Our platform combines proprietary vertical indexes with LLM-optimized retrieval systems to power AI agents, applications, and enterprise workflows. We are solving hard problems across search, large language models, and large-scale infrastructure to make AI systems more reliable, transparent, and useful.</p> <p>Our team includes engineers, researchers, product builders, and operators who care about solving meaningful problems and delivering real-world impact. Whether you are improving core infrastructure, shaping product experiences, or helping bring new AI capabilities to market, your work will help define how modern AI finds and uses knowledge.</p></div><h4 class="AnswerParser_AnswerParserH3__Mpe4s"><span style="text-decoration: underline;">About the Role</span></h4> <p>As a Site Reliability Engineer, you will own parts of the reliability, observability, and incident response posture for You.com’s production services. Your work will ensure that every user query, every API call, and every data pipeline runs with measurable, defensible uptime, and when something breaks, the tools and dashboards you developed will help the team identify the issue, respond, and learn from it.  Additionally, you will partner with teams to help them implement best practices, establish reliability objectives, and ensure the engineering team can build reliable services with minimal friction.</p> <h4 class="AnswerParser_AnswerParserH3__Mpe4s"><span style="text-decoration: underline;">Responsibilities</span></h4> <ul> <li><strong>Instrument services </strong>end-to-end using OpenTelemetry metrics and structured logging to ensure every critical path is measurable.</li> <li><strong>Develop and maintain SRE standards and patterns (instrumentation guidelines, incident playbooks, service templates) that engineering teams adopt by default in new and existing services.</strong><strong>Build internal tooling and automation in Python, Bash and Terraform to improve deployment safety, reliability, and operational efficiency.</strong></li> <li><strong>Design and maintain actionable dashboards</strong> that surface real user impact, not vanity metrics, for service owners and leadership.</li> <li><strong>Tune alerting rules continuously </strong>to maximize signal-to-noise ratio; tie alerts to SLO-based