Site Reliability Engineer Jobs in San Francisco, USA

73 verified site reliability engineer openings in San Francisco

  • Director of Site Reliability Engineering

    Stellar · San Francisco, USA

    On-site
    about 1 month agoApply →

    Director of Site Reliability Engineering at Stellar. Apply via Ashby.

  • Senior Site Reliability Engineer

    Carta · San Francisco, USA

    On-site
    about 1 month agoApply →

    <div class="content-intro"><h2>The Company You’ll Join</h2> <p>Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private equity and private credit. Trusted by 65,000+ companies in 160+ countries, Carta’s platform of software and services lays the groundwork so you can build, invest, and scale with confidence.</p> <p>Carta’s Fund Administration platform supports 9,000+ funds and SPVs, representing nearly $185B in assets under management, with tools designed to enhance the strategic impact of fund CFOs. Recognized by Fortune, Forbes, Fast Company, Inc. and Great Places to Work, Carta is shaping the future of private market infrastructure.</p> <p>Together, Carta is creating the end-to-end ERP platform for private markets. Traditional ERP solutions don’t work for Private Funds. Private capital markets need a comprehensive software solution to replace outdated spreadsheets and fragmented service providers. Carta’s software for the Office of the Fund CFO does just that - it’s a new category of software to make private markets look more like public markets - a connected ERP for private capital. </p> <p>For more information about our offices and culture, check out our <a href="https://carta.com/careers/">Carta careers page</a>.</p></div><h2><strong>The Problems You'll Solve</strong></h2> <p>At Carta, our employees set out on a mission to unlock the power of equity ownership for more people in more places. We believe that the problems we solve today unlock the opportunities of tomorrow.  As a <strong>Senior</strong> <strong>Site Reliability Engineer,</strong> you’ll work to: </p> <ul> <li>Build and scale our internal  platform offerings (compute, storage and networking services) to ensure the reliability, and performance of our applications.</li> <li>Design and implement monitoring, alerting, and incident response systems.</li> <li>Collaborate with application software engineers (as needed) to guide their design and ensure it scales for what Carta needs in the long run.</li> <li>Act as an agent of change and push boundaries to incrementally improve our systems as we expand globally.</li> </ul> <h2><strong>The Team You'll Work With</strong></h2> <p>You’ll be joining the Infrastructure Engineering team at Carta. The Infrastructure Engineering team is responsible for providing secure, reliable, scalable and performant infrastructure to Carta’s customers and developers.</p> <p>We are Software and Infrastructure Engineers who specialize in cloud computing, networking, systems design and architecture, storage, real time data telemetry, associated automation, tooling and proc

  • Senior Network & Site Reliability Engineer

    Alembic · San Francisco HQ, USA

    On-site
    about 1 month agoApply →

    Senior Network & Site Reliability Engineer at Alembic. Apply via Ashby.

  • Senior Manager, Site Reliability Engineering

    Tubi · San Francisco, USA

    Hybrid
    about 1 month agoApply →

    <p><strong>About the Role:</strong></p> <p>Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation.</p> <p>We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability Engineering team. You are more than a people manager or a tech lead; you are the strategic leader responsible for architecting our reliability roadmap. You will build and mentor a team of talented engineers, foster a culture of blameless learning and continuous improvement, and champion the engineering practices that allow us to balance rapid innovation with rock-solid stability. You will be a key influencer in our engineering leadership, partnering with peers across the organization to ensure reliability is a shared responsibility and a core tenet of our engineering culture.</p> <p><strong>What You'll Do:</strong></p> <ul> <li><strong>Team Leadership & Mentorship:</strong> <ul> <li>Lead, mentor, and grow a team of Site Reliability Engineers. Foster a culture of innovation and technical excellence where engineers feel empowered to do their best work. Provide personalized coaching, create professional development plans, and guide the careers of senior and emerging talent within the team.</li> <li>Establish equitable, sustainable on-call practices (including global coverage where applicable) that protect focus time and avoid burnout.</li> <li>Define team rituals - runbook reviews, game days, and incident retros - that reinforce quality and learning.</li> </ul> </li> <li><strong>Strategic Planning & Vision:</strong> Define and drive the multi-year technical strategy and vision for Tubi’s observability, and automation platforms. Partner with infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE roadmap with business objectives. Champion a data-driven approach to reliability, using Service Level Objectives (SLOs) and error budgets to facilitate productive conversations about risk and feature velocity.</li> <li><strong>Operational Excellence & Incident Management:</strong>  <ul> <li>Own the end-to-end availability, performance, and efficiency of our critical user-facing services. Evolve our incident response

  • Site Reliability Engineer

    Cognition · San Francisco, USA

    On-site
    about 1 month agoApply →

    Site Reliability Engineer at Cognition. Apply via Ashby.

  • Senior Site Reliability Engineer

    Alembic · San Francisco HQ, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer at Alembic. Apply via Ashby.

  • Sr. Site Reliability Engineer

    Prosper · San Francisco, CA

    On-site
    about 2 months agoApply →

    Your role in our mission   You will be a senior technical contributor on the SRE team, responsible for the reliability, scalability, and security of Prosper’s Cloud Platform portfolio. This is as much of a platform engineering role as it is SRE role — you will maintain the applications that run on our platform, drive alignment to platform standards, and ensure services stay current within the framework and dependency realm.We are building an agentic AI-first operations model where AI agents handle investigations, deployments, audits, and optimizations — and you will be at the center of designing and governing that system. You will share the ownership of application-layer reliability, CI/CD pipelines, and observability while simultaneously building the skills, rules, and guardrails that allow AI agents to operate safely alongside human engineers.

  • Senior Site Reliability Engineer

    Hive · San Francisco, San Francisco

    On-site
    about 2 months agoApply →

    About Hive Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted by hundreds of the world's largest and most innovative organizations. The company empowers developers with a portfolio of best-in-class, pre-trained AI models, serving billions of customer API requests every month. Hive also offers turnkey software applications powered by proprietary AI models and datasets, enabling breakthrough use cases across industries. Together, Hive’s solutions are transforming content moderation, brand protection, sponsorship measurement, context-based ad targeting, and more. Hive has raised over $120M in capital from leading investors, including General Catalyst, 8VC, Glynn Capital, Bain & Company, Visa Ventures, and others. We have over 250 employees globally in our San Francisco, Seattle, and Delhi offices. Please reach out if you are interested in joining the future of AI! DevOps and Systems Team Our unique machine learning needs led us to open our own data centers, with an emphasis on distributed high performance computing integrating GPUs. Even with these data centers, we maintain a hybrid infrastructure with public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able to thrive in an unstructured environment and takes automation seriously. You believe there is no task that can’t be automated and no server scale too large. You take pride in optimizing performance at scale in every part of the stack and never manually performing the same task twice.

  • Senior Site Reliability Engineer, Spend

    Airwallex · US - San Francisco, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer, Spend at Airwallex. Apply via Ashby.

  • Staff Site Reliability Engineer, Tech Lead

    Unify · San Francisco Office, USA

    On-site
    about 2 months agoApply →

    Staff Site Reliability Engineer, Tech Lead at Unify. Apply via Ashby.

  • Senior Staff Site Reliability Engineer, Tech Lead

    Unify · San Francisco Office, USA

    On-site
    about 2 months agoApply →

    Senior Staff Site Reliability Engineer, Tech Lead at Unify. Apply via Ashby.

  • Senior Site Reliability Engineer

    Unify · San Francisco Office, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer at Unify. Apply via Ashby.

  • Senior Site Reliability Engineer

    Hyperbolic · San Francisco, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer at Hyperbolic. Apply via Ashby.

  • Lead Site Reliability Engineer

    Stuut ai · San Francisco, USA

    On-site
    about 2 months agoApply →

    Lead Site Reliability Engineer at Stuut ai. Apply via Ashby.

  • Senior Site Reliability Engineer

    Anyscale · San Francisco or Palo Alto, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer at Anyscale. Apply via Ashby.

  • Site Reliability Engineer (SRE)

    Airapps · San Francisco, USA

    On-site
    about 2 months agoApply →

    Site Reliability Engineer (SRE) at Airapps. Apply via Ashby.

  • Senior Site Reliability Engineer

    youcom · San Francisco (Hybrid), San Francisco (Hybrid)

    Hybrid
    about 2 months agoApply →
  • Site Reliability Engineer

    vantagescore · San Francisco, CA

    On-site
    about 2 months agoApply →
  • Senior Manager, Site Reliability Engineering

    tubitv · San Francisco, CA (Hybrid)

    Hybrid
    about 2 months agoApply →
  • Site Reliability Engineer (SRE)

    thinkingmachines · San Francisco, San Francisco

    On-site
    about 2 months agoApply →