Site Reliability Engineer
SingleStore · Seattle, USA
On-site12 days agoApply →<p><strong>Position Overview</strong></p> <p>SingleStore is seeking a Site Reliability Engineer to help optimize and scale our managed service offering across all three major cloud providers. In this role, you will be at the intersection of leading technology trends – A highly performant distributed database, managed by Kubernetes, running in the cloud.&nbsp; This is a great opportunity to push the boundaries with a cloud-focused SRE role.&nbsp;&nbsp;</p> <p>This is a development role, requiring an engineering mindset to solve operational challenges.&nbsp; You will be part of a globally distributed team of engineers, helping to drive SRE practices across the company.&nbsp; Through infrastructure automation, you will help us grow our service across multiple cloud platforms.&nbsp; This requires a relentless focus on eliminating manual processes.&nbsp; You will also leverage our monitoring platform to improve the overall customer experience by systematically identifying and fixing any issues impacting our customers.&nbsp; As an SRE, you will also help diagnose issues on the platform, leveraging a deep understanding of the SingleStore query engine along with the backend infrastructure.&nbsp;&nbsp;</p> <p><strong>Roles and Responsibilities</strong></p> <ul> <li>Develop automation platform to manage infrastructure rollouts across cloud providers</li> <li>Optimize telemetry platform to identify customer impacting events while providing relevant data to drive debugging</li> <li>Partner with engineering team to optimize performance of services for cloud architecture</li> <li>Debug Live Site events and conduct follow-up postmortem and RCA analysis</li> <li>Participate in an SLA-driven on-call rotation, which will include after-hours, weekend, and rotating holiday participation.</li> </ul> <p><strong>Required Skills and Experience</strong></p> <ul> <li>0 - 2 years of demonstrated experience working as a Site Reliability Engineer.&nbsp; Recent graduates encouraged to apply.&nbsp;</li> <li>Infrastructure automation experience. Scripting experience (Python, Bash) a plus.</li> <li>Experience with the Prometheus monitoring stack. Experience with Grafana, Mimir and Loki is a plus.</li> <li>Knowledge of Kubernetes and the container ecosystem</li> <li>Strong cross group collaboration and communication skills</li> <li>Familiar with at least one of AWS, Azure, or Google Cloud</li> <li>Experience debugging, diagnosing and troubleshooting complex, production software</li> <li>B.S. Degree in Computer Science or related field</li> </ul> <hr> <p>SingleStore is a global database company that empowers the world’s leading organizations to build
Senior Site Reliability Engineer I
axon · Seattle, United States
On-site21 days agoApply →Site Reliability Engineer (Senior or Staff), Infrastructure Security
MongoDB · Austin; New York City; San Francisco; Seattle; United States, USA
On-siteabout 1 month agoApply →<p>We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.</p> <p>The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture.</p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation.</p> <p>This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones.</p> <h3>Responsibilities:</h3> <p>Cloud Security Design and Implementation:&nbsp;&nbsp;</p> <ul> <li>Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM)</li> </ul> <p>Automation and Monitoring:&nbsp;&nbsp;</p> <ul> <li>Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection</li> </ul> <p>Security Tooling:</p> <ul> <li>Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM&nbsp;</li> </ul> <h3>Qualifications:</h3> <p>Experience:&nbsp;&nbsp;</p> <ul> <li>6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role</li> </ul> <p>Security Mindset:</p> <ul> <li>A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking layers to cloud provider configurations. Proven experience in leading projects within security-focused areas, such as runtime scanning, security observability, CSPM, and more</li> </ul> <p>Cloud Expertise:&nbsp;&nbsp;</p> <ul> <li>Strong experience with at least one cloud platform (AWS, Azure, GCP), including expertise in IAM, VPC networking, security groups, and cloud security tools (e.g., GuardDuty, Security Hub, CloudTrail)</li> </ul> <p>Coding/Automation:&nbsp;&nbsp;</p> <ul>
Senior Site Reliability Engineer, Fleet Management
MongoDB · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, USA
On-siteabout 1 month agoApply →<h3>The Team</h3> <p>Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems.</p> <p>The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model.</p> <p>This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin.</p> <h3>Responsibilities</h3> <ul> <li>Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB</li> <li>Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems</li> <li>Participate in a 24/7 on-call rotation to resolve critical issues</li> <ul> <li>Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice</li> </ul> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 6+ years of experience in software development and operating distributed systems</li> <li>Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests)</li> <li>Have deep experience using and extending containerization technologies, preferably Kubernetes</li> <li>Have a solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS)</li> <li>Possess a customer focused mindset, treating internal developers as your primary users</li> <li>Have strong operational ownership, including a track record of debugging complex production issues and driving them to resolution</li> <li>Prefer automation over manual processes ("allergic to ops work")</li> <li>We are a small team of software engineers with a strong bias toward building software solutions to eliminate toil</li> </ul> <h3>Strong candidates may also have experience w
Senior Site Reliability Engineer I
Axon · Seattle, USA
On-siteabout 1 month agoApply →<div class="content-intro"><h4><strong>Join Axon and be a Force for Good.</strong></h4> <p>At Axon, we’re on a mission to Protect Life. We’re explorers, pursuing society’s most critical safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with candor and care, seeking out diverse perspectives from our customers, communities and each other.<br><br>Life at Axon is fast-paced, challenging and meaningful. Here, you’ll take ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.</p></div><h3><strong>Your Impact</strong></h3> <p>Are you an engineer who gets excited about the challenge of making complex distributed systems observable — not just instrumenting them, but designing the infrastructure that makes traces, metrics, and logs useful at scale? In this role, you will help build and evolve Axon's next-generation observability platform, enabling the entire engineering organization to understand and operate their services with confidence.<br><br>You'll work across the full observability stack: from distributed tracing adoption (OpenTelemetry, Jaeger) to log infrastructure (Loki, Alloy) to metrics (Cortex, Prometheus, Grafana). You'll partner directly with Axon's engineering teams to drive adoption of modern observability practices and build the tooling that makes our platform self-service for the teams that depend on it.<br><br>You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics, logging, tracing, and alerting infrastructure across dozens of environments globally.</p> <p>The ideal candidate has a strong infrastructure engineering background, is comfortable working across cloud-native systems, and cares about both the technical depth and the developer experience of observability. You'll thrive here if you have opinions about what good observability looks like, and enjoy the challenge of making it real in a large, fast-moving organization.</p> <p><strong>Location:&nbsp;</strong>This role is based out of our Seattle office and follows a hybrid schedule. We rely on in-person collaboration and ask that team members work onsite Tuesdays through Fridays, with the flexibility to work remotely on Mondays, unless there is an approved workplace accommodation. We believe that connection fuels innovation, and our in-office culture is designed to foster meaningful teamwork, mentorship, and shared success.</p> <h3><strong>What You'll Do</strong></h3> <ul> <li>Own and evolve Axon's distributed tracing infrastructure, including Jaeger and OpenTelemetry-based instrumentation, driving adoption across Axon
Senior Site Reliability Engineer I
Axon · Seattle, USA
On-siteabout 1 month agoApply →<div class="content-intro"><h4><strong>Join Axon and be a Force for Good.</strong></h4> <p>At Axon, we’re on a mission to Protect Life. We’re explorers, pursuing society’s most critical safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with candor and care, seeking out diverse perspectives from our customers, communities and each other.<br><br>Life at Axon is fast-paced, challenging and meaningful. Here, you’ll take ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.</p></div><h3>Your Impact</h3> <p>As a Senior SRE on the APX SRE CloudOps team, you will design and build the cloud infrastructure and automation platforms that Axon's product engineering teams depend on. You will architect solutions for multi-cloud environments (Azure, AWS), FedRAMP compliance boundaries, and large-scale Kubernetes platforms serving production workloads across multiple regions. You will spend a significant portion of your time writing code: building services, APIs, and internal tooling using languages like Go and Python. You will also participate in on-call rotations and incident response, using operational experience to drive reliability improvements and inform platform investment. This role combines software engineering depth with cloud architecture at scale and production ownership. <br><br><strong>Location -</strong> This role is based out of our Atlanta, Seattle, or Boston office and follows a hybrid schedule. We rely on in-person collaboration and ask that team members work onsite Tuesdays through Fridays, with the flexibility to work remotely on Mondays, unless there is an approved workplace accommodation. We believe that connection fuels innovation, and our in-office culture is designed to foster meaningful teamwork, mentorship, and shared success.</p> <h3>What You'll Do</h3> <ul> <li>Design and implement services, APIs, and automation tooling that improve cloud infrastructure reliability, scalability, and operational efficiency.</li> <li>Write production-quality code in Go, Python, or similar. Lead design reviews and set code quality standards within the team.</li> <li>Build robust, easy-to-use foundational platforms and tools that enable engineering teams to provision services rapidly, consistently, securely, and cost-effectively.</li> <li>Architect cloud infrastructure across Azure and AWS, including Kubernetes platforms, networking, and identity boundaries.</li> <li>Participate in on-call rotations and incident response. Use operational experience to identify systemic reliability risks and drive platform improvements.</li> <li>Foresee and mitigate risks in infrastructure design: single points of fai
Senior Site Reliability Engineer
Hive · Seattle, Seattle
On-siteabout 2 months agoApply →About Hive Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted by hundreds of the world's largest and most innovative organizations. The company empowers developers with a portfolio of best-in-class, pre-trained AI models, serving billions of customer API requests every month. Hive also offers turnkey software applications powered by proprietary AI models and datasets, enabling breakthrough use cases across industries. Together, Hive’s solutions are transforming content moderation, brand protection, sponsorship measurement, context-based ad targeting, and more. Hive has raised over $120M in capital from leading investors, including General Catalyst, 8VC, Glynn Capital, Bain & Company, Visa Ventures, and others. We have over 250 employees globally in our San Francisco, Seattle, and Delhi offices. Please reach out if you are interested in joining the future of AI! DevOps and Systems Team Our unique machine learning needs led us to open our own data centers, with an emphasis on distributed high performance computing integrating GPUs. Even with these data centers, we maintain a hybrid infrastructure with public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able to thrive in an unstructured environment and takes automation seriously. You believe there is no task that can’t be automated and no server scale too large. You take pride in optimizing performance at scale in every part of the stack and never manually performing the same task twice.
Site Reliability Engineer
singlestore · Seattle, Seattle
On-siteabout 2 months agoApply →Sr. Site Reliability Engineer
pitchbookdata · Seattle, United States
On-siteabout 2 months agoApply →Sr. Site Reliability Engineer
pitchbookdata · Seattle, United States
On-siteabout 2 months agoApply →Site Reliability Engineer (Senior or Staff), Infrastructure Security
mongodb · Austin; New York City; San Francisco; Seattle; United States, Austin; New York City; San Francisco; Seattle; United States
On-siteabout 2 months agoApply →Senior Site Reliability Engineer, Fleet Management
mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States
On-siteabout 2 months agoApply →Manager, Site Reliability Engineering - Fleet Management
mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle
On-siteabout 2 months agoApply →Senior Site Reliability Engineer I
axon · Seattle, United States
On-siteabout 2 months agoApply →Senior Site Reliability Engineer I
axon · Seattle, United States
On-siteabout 2 months agoApply →Site Reliability Engineer, Discovery
andurilindustries · Seattle, United States
On-siteabout 2 months agoApply →Senior Site Reliability Engineer, Production Engineering
andurilindustries · Seattle, United States
On-siteabout 2 months agoApply →