Site Reliability Engineer Jobs in Austin, USA

23 verified site reliability engineer openings in Austin

  • Senior Site Reliability Engineer, Database Infrastructure

    zello · Austin, Texas, USA

    On-site
    17 days agoApply →
  • Senior Site Reliability Engineer - Developer Experience

    Dimensional · Austin, USA

    On-site
    19 days agoApply →
  • LEAD SITE RELIABILITY ENGINEER

    Cox · Austin TX, USA

    On-site
    19 days agoApply →
  • Senior Site Reliability Engineer

    2k · Austin, United States

    On-site
    24 days agoApply →
  • Site Reliability Engineer (Senior or Staff), Infrastructure Security

    MongoDB · Austin; New York City; San Francisco; Seattle; United States, USA

    On-site
    about 1 month agoApply →

    <p>We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.</p> <p>The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture.</p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation.</p> <p>This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones.</p> <h3>Responsibilities:</h3> <p>Cloud Security Design and Implementation:  </p> <ul> <li>Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM)</li> </ul> <p>Automation and Monitoring:  </p> <ul> <li>Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection</li> </ul> <p>Security Tooling:</p> <ul> <li>Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM </li> </ul> <h3>Qualifications:</h3> <p>Experience:  </p> <ul> <li>6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role</li> </ul> <p>Security Mindset:</p> <ul> <li>A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking layers to cloud provider configurations. Proven experience in leading projects within security-focused areas, such as runtime scanning, security observability, CSPM, and more</li> </ul> <p>Cloud Expertise:  </p> <ul> <li>Strong experience with at least one cloud platform (AWS, Azure, GCP), including expertise in IAM, VPC networking, security groups, and cloud security tools (e.g., GuardDuty, Security Hub, CloudTrail)</li> </ul> <p>Coding/Automation:  </p> <ul>

  • Site Reliability Engineer (Senior or Staff), Atlas

    MongoDB · Austin; Boston; Chicago; Miami; New York City; Philadelphia; Pittsburgh; Raleigh; United States; Was, USA

    On-site
    about 1 month agoApply →

    <h3>The Team</h3> <p>This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. </p> <p>The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. </p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. </p> <h3>Role Overview</h3> <p>We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. </p> <h3>The ideal candidate should</h3> <ul> <li>Have 5+ years of experience running critical systems at scale</li> <li>Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”)</li> <li>Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment</li> <li>A strong understanding of how to run a large scale Linux environment, including low level fundamentals </li> <li>Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python)</li> <li>Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc)</li> </ul> <h3>Special Requirements:</h3> <ul> <li>Be a US Citizen</li> </ul> <h3>Expectations</h3> <ul> <li>Participate in the development of a reliable and resilient multi-cloud platform that hosts business critical applications for a wide & varied range of customer applications</li> <li>Collaborate with service-owning teams to provide internal support, solve technical challenges and adapt or build tooling to solve novel use cases in a generic fashion</li> <li>Participate in a 24/7 on-call rotation to swiftly resolve issues related to any disruption of our customer facing Atlas fleet, ensuring minimal disruption and high availability</li> </ul> <h3>About MongoDB</h3> <p>MongoDB is built for change, empowering our customers

  • Senior Site Reliability Engineer, Fleet Management

    MongoDB · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, USA

    On-site
    about 1 month agoApply →

    <h3>The Team</h3> <p>Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems.</p> <p>The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model.</p> <p>This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin.</p> <h3>Responsibilities</h3> <ul> <li>Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB</li> <li>Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems</li> <li>Participate in a 24/7 on-call rotation to resolve critical issues</li> <ul> <li>Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice</li> </ul> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 6+ years of experience in software development and operating distributed systems</li> <li>Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests)</li> <li>Have deep experience using and extending containerization technologies, preferably Kubernetes</li> <li>Have a solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS)</li> <li>Possess a customer focused mindset, treating internal developers as your primary users</li> <li>Have strong operational ownership, including a track record of debugging complex production issues and driving them to resolution</li> <li>Prefer automation over manual processes ("allergic to ops work")</li> <li>We are a small team of software engineers with a strong bias toward building software solutions to eliminate toil</li> </ul> <h3>Strong candidates may also have experience w

  • Sr Site Reliability Engineer

    Realtor.com Careers · Austin, USA

    On-site
    about 1 month agoApply →

    <div class="content-intro"><p>Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate for over 25 years, connecting buyers, sellers, and renters with trusted insights and expert guidance to find their perfect home. Through its robust suite of tools, Realtor.com® not only makes a significant impact on the real estate industry at large, but for consumers, navigating the biggest purchase they will make in their life, by providing a user experience that is easy to use, easy to understand, and most of all, easy to make decisions.</p> <p>Join us on our mission to empower more people to find their way home by breaking barriers to entry, making the right connections, and building confidence through expert guidance.</p></div><p>We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization, reporting to the Director, Operations Excellence. This role will contribute to the reliability, observability, and operational excellence of our platform infrastructure serving millions of users. As a Senior SRE, you will be a strong technical contributor who implements best practices, solves complex problems, and enables our 600+ engineers to deliver exceptional customer experiences. You will work on critical platform systems including EKS infrastructure, Skyway (CI/CD), Frontdoor (Tyk API Gateway), Pantheon (Apollo GraphQL Federation), and our observability stack, while contributing to chaos engineering practices and cost optimization initiatives with measurable ROI.</p> <p>What You'll Do:</p> <p>Platform Reliability & Infrastructure</p> <ul> <li>Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures</li> <li>Support reliability of critical services: Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure</li> <li>Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency</li> <li>Implement reliability patterns including circuit breakers, graceful degradation, and automated failover</li> </ul> <p>Observability & Cost Optimization</p> <ul> <li>Implement observability solutions using NewRelic for APM, distributed tracing, metrics, and logging for rapid troubleshooting</li> <li>Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams</li> <li>Identify infrastructure cost optimization opportunities and implement FinOps practices including rightsizing and resource lifecycle management</li> <li>Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD)</li> </ul>

  • Sr Site Reliability Engineer

    rdccareers · Austin, United States

    On-site
    about 1 month agoApply →
  • Site Reliability Engineer, Metal

    Tenstorrent · Austin, USA

    On-site
    about 2 months agoApply →

    <div class="content-intro"><p>Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.</p></div><p>Tenstorrent is building large-scale AI systems across internal clusters and customer deployments. This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and production-ready.</p> <p>This role is<strong> </strong>hybrid, based out of Toronto, ON; Austin, TX; or Santa Clara, CA.</p> <p>We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.</p> <p> </p> <p><strong>Who You Are</strong></p> <ul> <li data-start="427" data-end="528">Experienced in site reliability, infrastructure, or systems engineering in distributed environments.</li> <li data-start="427" data-end="528">Strong Linux systems knowledge with the ability to troubleshoot complex multi-layer issues.</li> <li data-start="427" data-end="528">Proficient with observability tools such as Prometheus, Grafana, and alerting systems.</li> <li data-start="427" data-end="528">Comfortable with scripting and automation using Python, Go, or similar languages.</li> <li data-start="427" data-end="528">Solid understanding of networking fundamentals and how systems behave at scale.</li> </ul> <p> </p> <p><strong>What We Need</strong></p> <ul> <li data-start="907" data-end="1015">Ensure reliability and operational health of Tenstorrent systems across internal and customer environments.</li> <li data-start="907" data-end="1015">Troubleshoot complex issues across compute, networking, and software layers.</li> <li data-start="907" data-end="1015">Partner with engineering teams and customers to resolve production incidents.</li> <li data-start="907" data-end="1015">Design and improve monitoring, observability, a

  • ​Senior Site Reliability Engineer

    Cox · Austin TX,

    On-site
    about 2 months agoApply →
  • Senior Site Reliability Engineer

    Zello · Austin, USA

    On-site
    about 2 months agoApply →

    Senior Site Reliability Engineer at Zello. Apply via Ashby.

  • Site Reliability Engineer II

    Restaurant365 · Austin, CA

    On-site
    about 2 months agoApply →

    Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365’s culture is focused on empowering team members to produce top-notch results while elevating their skills. We’re constantly evolving and improving to make sure we are and always will be “Best in Class” ... and we want that for you too! This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH.    The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications. Qualified candidates will demonstrate growing expertise in site reliability practices, with skills in incident response, system monitoring, automation, and performance troubleshooting. You will collaborate with DevOps, development, and infrastructure teams to resolve moderately complex issues, propose improvements, and strengthen the reliability, scalability, and security of our SaaS platform. 

  • Senior Site Reliability Engineer

    securityscorecard · Austin, TX (Hybrid)

    Hybrid
    about 2 months agoApply →
  • Senior Site Reliability Engineer

    ujet · Austin, US

    Remote
    about 2 months agoApply →
  • Senior Site Reliability Engineer

    tripactions · Austin, TX

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer, Metal

    tenstorrent · Austin, Canada

    On-site
    about 2 months agoApply →
  • Senior Site Reliability Engineer, Fleet Management

    mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States

    On-site
    about 2 months agoApply →
  • Manager, Site Reliability Engineering - Fleet Management

    mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer (Senior or Staff), Infrastructure Security

    mongodb · Austin; New York City; San Francisco; Seattle; United States, Austin; New York City; San Francisco; Seattle; United States

    On-site
    about 2 months agoApply →