Staff Site Reliability Engineer - Site Experience
Reddit · Dublin, Ireland
On-site2 days agoApply →<div class="content-intro"><div class="c-message_kit__blocks c-message_kit__blocks--rich_text"> <div class="c-message__message_blocks c-message__message_blocks--rich_text" data-qa="message-text"> <div class="p-block_kit_renderer" data-qa="block-kit-renderer"> <div class="p-block_kit_renderer__block_wrapper p-block_kit_renderer__block_wrapper--first"> <div class="p-rich_text_block"> <div class="p-rich_text_section">Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit <a class="c-link" href="http://www.redditinc.com/" target="_blank" data-stringify-link="http://redditinc.com" data-sk="tooltip_parent">www.redditinc.com</a>.</div> </div> </div> </div> </div> </div></div><p>As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across web, mobile, APIs, feeds, media delivery, and real time systems is fast, reliable, and resilient.</p> <p>We are looking for a Staff Site Reliability Engineer to lead reliability engineering initiatives for critical user facing systems at internet scale. In this role, you will partner closely with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across Reddit’s most business critical experiences.</p> <p>This is a highly technical leadership role for someone who thrives in large-scale distributed systems, enjoys solving complex reliability challenges, and can influence engineering culture across the organization.</p> <h4><strong>What you’ll do:</strong></h4> <ul> <li><strong>Lead Reliability Engineering for User Experience</strong></li> <ul> <li>Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences.</li> </ul> <li><strong>Architect for Scale</strong></li> <ul> <li>Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failov
Staff Technical Program Manager, Site Reliability Engineering
MongoDB · Cork; Dublin; Ireland, Ireland
On-siteabout 1 month agoApply →<p>As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale.&nbsp;</p> <p>This role can be based out of our Dublin or Cork office or remotely in Ireland.</p> <h3>What You'll Do</h3> <ul> <li>Drive Program Planning &amp; Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time</li> <li>Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement</li> <li>Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts</li> <li>Build Scalable Systems &amp; Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently</li> </ul> <h3>Requirements</h3> <ul> <li>8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams</li> <li>Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change</li> <li>Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs)</li> <li>Skilled at shaping roadmaps and managing dependencies</li> <li>Able to query and interpret metrics, logs, or other data sources to inform decisions and communicate risk</li> <li>Excellent communicator—clear, concise, and calm—across engineers, cross-functional partners, and executives</li> <li>Low-ego, highly collaborative, and motivated by ownership of hard problems end to end</li> </ul> <h3>Nice to Have</h3> <ul> <li>Hands-on or close-partner experience with Kubernetes, cloud networking, or observability stacks (metrics, logs, tracing, alerting)</li> <li>Prior experience working with or alongside SRE teams</li> <li>Background in large-scale cloud infrastructure or platform e
Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
MongoDB · Cork; Dublin; Ireland, Ireland
On-siteabout 1 month agoApply →<h3>The Team</h3> <p>MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently.</p> <p>You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture.</p> <p>This role can be based out of either our Dublin or Cork office or remotely in Ireland.</p> <h3>The ideal candidate should:&nbsp;</h3> <ul> <li>Have 6+ years of experience working on software development and operating distributed systems</li> <li>Proficiency in Python, Go, or a similar language</li> <li>Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs.</li> <li>Possess a customer-focused mindset</li> <li>Value efficiency in processes and operations</li> <li>Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil</li> <li>Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market</li> <li>Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure</li> <li>Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)</li> </ul> <h3>Responsibilities:&nbsp;</h3> <ul> <li>Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs</li> <li>Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing</li> <li>Identify and configure key metrics to detect incidents and quantify service health, availability, and performance</li> <li>Participate in a 24/7 on-call rotation to resolve issues involving the storage infrastructure</li> <li>Become an expert in infrastructure performance, helping us optimize from the application level all the way to the kernel</li> </ul> <h3>Strong candidates may also have experience with:</h3> <ul> <li>Leading major architectural shifts, such as movi
Site Reliability Engineer 3
MongoDB · Dublin, Ireland
On-siteabout 1 month agoApply →<p>MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications.</p> <p>The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable boundary between their areas of responsibility.</p> <p>We are looking to speak to candidates who are based in Dublin for our hybrid working model.</p> <h3>Responsibilities</h3> <ul> <li>Design and build the infrastructure for a global cloud service that comprises hundreds of thousands of MongoDB clusters, processes a billion metrics per day, and replicates tens of billions of database writes to our backup service</li> <li>Design, implement, and troubleshoot the automation and monitoring of services that seamlessly spans the globe - including several cloud providers</li> <li>Become an expert in infrastructure performance, helping us optimize from the application level all the way through the firmware</li> <li>Build for resilience. Our goal is that nobody’s pager goes off, ever. Are we there yet? No. Are we really close? Very. While we work on that - participate in a weekly on-call rotation</li> <li>Improve our infrastructure capabilities, optimizing for cost, simplicity, and maintainability</li> </ul> <h3>Requirements</h3> <ul> <li>3+ years of experience running a mission critical service at scale in a Linux environment</li> <li>Firm grasp of at least one modern programming language, beyond basic scri
Manager, Site Reliability Engineering - Storage Layer Service
MongoDB · Dublin, Ireland
On-siteabout 1 month agoApply →<p>MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently.</p> <p>As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture.</p> <p>We are looking to speak to candidates who are based in Dublin for our hybrid working model.</p> <h3>Responsibilities</h3> <ul> <li>Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers</li> <li>Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs</li> <li>Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges</li> <li>Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations</li> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams</li> <li>Possess a customer-focused mindset, treating internal developers as your primary users</li> <li>Value efficiency in processes and operations, and have a track record of optimizing team workflows</li> <li>Prefer automation over manual processes, fostering a culture of building software solutions to eliminate toil</li> <li>Have deep technical familiarity with Kubernetes ecosystems, containerization technologies, and modern IaC tooling (e.g., Terraform, Crossplane, or Operators) so you can effectively guide the team's technical decisions</li> <li>Have operated or supported stateful storage or database systems at scale and are comfortable with durability, consistency and recovery trade-offs</li> <li>Excel at translating complex business and engineering requirements into actionable, phased technical roadmaps</li> <li>Have a high level of empathy, responsibi
Site Reliability Engineer
StepStone Group · Dublin, Ireland
On-siteabout 1 month agoApply →<div class="content-intro"><p><span data-teams="true"><span class="ui-provider a b c d e f g h i j k l m n o p q r s t u v w x y z ab ac ae af ag ah ai aj ak">We are global private markets specialists&nbsp;delivering tailored investment solutions,&nbsp;advisory services, and impactful, data driven&nbsp;insights to the world’s investors.&nbsp;Leveraging the power of our platform and&nbsp;our peerless intelligence across sectors,&nbsp;strategies, and geographies, we help&nbsp;identify the advantages and the answers&nbsp;our clients need to succeed.</span></span></p></div><p>The Site Reliability Engineer is responsible for designing, deploying and managing enterprise solutions utilizing various network, endpoint and cloud technologies. The role will provide subject matter expertise on complex cloud native technologies, topics and issues.</p> <p><strong>Responsibilities</strong></p> <ul> <li>Design, build, and maintain scalable cloud infrastructure on AWS, GCP, or Azure, with a focus on high availability and fault tolerance.</li> <li>Collaborate with software engineers to embed reliability best practices into the software development lifecycle.</li> <li>Participate in on-call rotations, lead incident response, and conduct thorough post-mortems to prevent recurrence.</li> <li>Develop and maintain infrastructure-as-code (IaC) using tools such as Terraform, and CloudFormation</li> <li>Optimize cloud resource utilization and cost management across multi-cloud or hybrid environments.</li> <li>Contribute to the design and improvement of CI/CD pipelines and deployment automation.</li> <li>Ensure cloud environments adhere to financial industry security and compliance standards.</li> <li>Document systems, runbooks, and processes to support team knowledge sharing.</li> </ul> <p><strong>Required Qualifications</strong></p> <ul> <li>3–5 years of experience in a Site Reliability Engineering, DevOps, or Cloud Infrastructure role.</li> <li>Hands-on experience with one or more major cloud providers: AWS, GCP, or Azure.</li> <li>Proficiency in at least one scripting or programming language (Python, Go, Bash, etc.).</li> <li>Experience with infrastructure-as-code tools (Terraform, CloudFormation, Bicep).</li> <li>Strong understanding of networking fundamentals (DNS, TCP/IP, load balancing, VPNs).</li> <li>Familiarity with containerization and orchestration technologies (Docker, Kubernetes).</li> <li>Experience with observability tools such as Datadog, Prometheus, Grafana, or equivalent.</li> <li>Solid understanding of Linux/Unix systems administration.</li> </ul> <p><strong>Pre
Senior Site Reliability Engineer (R-19383)
Dnb · Dublin - Ireland, Dublin - Ireland
On-siteabout 2 months agoApply →The Senior Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, performance, and operability of production systems across our platforms, by applying software engineering practices to operations, with a focus on automation, observability, and incident response.
Site Reliability Engineer (SRE)
Airapps · Dublin, Ireland
On-siteabout 2 months agoApply →Site Reliability Engineer (SRE) at Airapps. Apply via Ashby.
Staff Site Reliability Engineer - Site Experience
reddit · Dublin, Ireland
On-siteabout 2 months agoApply →