Senior/Staff Site Reliability Engineer
Sage · New York, USA
On-site22 days agoApply →<p><strong>About Us</strong></p> <p>Sage is on a mission to improve care and quality of life for older adults, starting with those residing in senior living facilities. Falls are the leading cause of injury-related death among adults over 65. And yet, fall prevention and emergency response systems for older adults are archaic and ineffective. At Sage we've built a more modern way of understanding when older adults need help, including methods for residents to alert caregivers when in need of help, and corresponding software for caregivers to triage response. Our company mission is to create a product that our client counterparts love, and this role is a key part of that objective.</p> <p>Sage is a small, tight team of ambitious, multi-disciplinary entrepreneurs. We are a software-enabled, mission-driven company, and are focused only on the problems that are central to achieving that mission. At Sage, we work hard and fast but also know that to build a truly important company, we need to treat our work as a marathon, and not a sprint. The journey matters.</p> <p><strong>About this Role</strong></p> <p>Sage provides life-saving functionality that improves the lives of our older population. This role is critical to ensure Sage can live up to its mission to be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you’ll partner with engineering teams across the organization to achieve four 9s of uptime for our platform.</p> <p><strong>Responsibilities</strong></p> <ul> <li><strong>Design and evolve highly reliable system architectures</strong>, ensuring high availability, fault tolerance, and scalability across Sage’s production infrastructure.</li> <li><strong>Lead complex incident response efforts</strong>, coordinating across engineering teams to quickly diagnose and resolve production issues while driving thorough post-incident reviews and long-term reliability improvements.</li> <li><strong>Define and implement organization-wide observability practices</strong>, including metrics, logging, tracing, and actionable alerting to ensure strong visibility into system health.</li> <li><strong>Establish and maintain reliability standards</strong>, including defining SLIs, SLOs, and error budgets, and partnering with engineering teams to integrate these practices into the software development lifecycle.</li> <li><strong>Drive automation and infrastructure improvements</strong> that reduce operational toil and improve the efficiency and reliability of deployments, monitoring, and operational workflows.</li> <li><strong>Partner with engineering teams on system design and architecture reviews</strong>, ensuring reliability, scalability, and operational best practices are conside
Senior Software Engineer, Site Reliability Engineering
ridgeline · New York, NV
On-site24 days agoApply →Technology, DevOps/Site Reliability Engineer
btig27 · New York, San Francisco
On-site24 days agoApply →Staff Site Reliability Engineer
Tabs · New York City, USA
On-siteabout 1 month agoApply →Staff Site Reliability Engineer at Tabs. Apply via Ashby.
Staff Technical Program Manager, Site Reliability Engineering
MongoDB · Atlanta; Boston; Florida; Georgia; Maine; Miami; New Hampshire; New Jersey; New York; New York City;, USA
On-siteabout 1 month agoApply →<p>As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale.&nbsp;</p> <p>This role can be based remotely on the East Coast</p> <h3>What You'll Do</h3> <ul> <li>Drive Program Planning &amp; Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time</li> <li>Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement</li> <li>Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts</li> <li>Build Scalable Systems &amp; Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently</li> </ul> <h3>Requirements</h3> <ul> <li>8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams</li> <li>Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change</li> <li>Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs)</li> <li>Skilled at shaping roadmaps and managing dependencies</li> <li>Able to query and interpret metrics, logs, or other data sources to inform decisions and communicate risk</li> <li>Excellent communicator—clear, concise, and calm—across engineers, cross-functional partners, and executives</li> <li>Low-ego, highly collaborative, and motivated by ownership of hard problems end to end</li> </ul> <h3>Nice to Have</h3> <ul> <li>Hands-on or close-partner experience with Kubernetes, cloud networking, or observability stacks (metrics, logs, tracing, alerting)</li> <li>Prior experience working with or alongside SRE teams</li> <li>Background in large-scale cloud infrastructure or platform engineering</li> <li&g
Site Reliability Engineer (Senior or Staff)
MongoDB · Boston; Miami; New Jersey; New York City; Princeton; Raleigh; Washington DC, USA
On-siteabout 1 month agoApply →<p>Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems.</p> <p>The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams.</p> <p>We are looking to speak to candidates who can work East Coast hours.</p> <h3>The ideal candidate should</h3> <ul> <li>Have 6+ years of experience in software development and operating distributed systems</li> <li>Proficiency in Python, Go, or a similar language</li> <li>Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines</li> <li>Possess a customer-focused mindset</li> <li>Value efficiency in processes and operations</li> <li>Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil</li> <li>Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market</li> <li>Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure</li> <li>Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)</li> </ul> <h3>Expectations</h3> <ul> <li>Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products</li> <ul> <li>This includes, but is not limited to, contributing to open-source projects, or engineering software-based approaches like Kubernetes operators to streamline processes</li> </ul> <li>Own the onboarding flow other engineering teams follow when launching a new product or service</li> <li>Collaborate with other teams within Platform Engineering to ensure a consistent service-onboarding experience</li> <li>Provide internal support for our deployment systems, including answering questions and addressing issues</li> <li>Participate in a 24/7 on-call rotation to resolve issues involving the deployment infrastructure</li> </ul> <h3>About MongoDB</h3> <p>Mong
Site Reliability Engineer 3
MongoDB · New York City, USA
On-siteabout 1 month agoApply →<p>The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable boundary between their areas of responsibility.</p> <p>We are looking to speak to candidates who are based in New York City for our hybrid working model.</p> <h3>Responsibilities</h3> <ul> <li>Design and build the infrastructure for a global cloud service that comprises hundreds of thousands of MongoDB clusters, processes a billion metrics per day, and replicates tens of billions of database writes to our backup service</li> <li>Design, implement, and troubleshoot the automation and monitoring of services that seamlessly spans the globe - including several cloud providers</li> <li>Become an expert in infrastructure performance, helping us optimize from the application level all the way through the firmware</li> <li>Build for resilience. Our goal is that nobody’s pager goes off, ever. Are we there yet? No. Are we really close? Very. While we work on that - participate in a weekly on-call rotation</li> <li>Improve our infrastructure capabilities, optimizing for cost, simplicity, and maintainability</li> </ul> <h3>Requirements</h3> <ul> <li>3+ years of experience running a mission critical service at scale in a Linux environment</li> <li>Firm grasp of at least one modern programming language, beyond basic scripting</li> <li>Familiarity with web and network protocols and standards (HTTP, TLS, DNS, etc)</li> <li>Bachelor’s degree in Computer Science or equivalent experience</li> <li>Experience writing automation tools &amp; eagerness to "automate all the things"</li> </ul> <h3>Nice to have</h3> <ul> <li>Experience building large applications from scratch, complete with CI/CD infrastructure</li> <li>Experience in networking, security, hardware or OS performance tuning</li> <li>Experience with at least one of the major cloud providers (Amazon Web Services, Google Compute, Microsoft Azure)</li> <li>Experience managing kubernetes clusters or some other container orchestration infrastructure</li> <li>Experience with o
Site Reliability Engineer (Senior or Staff), Atlas
MongoDB · Austin; Boston; Chicago; Miami; New York City; Philadelphia; Pittsburgh; Raleigh; United States; Was, USA
On-siteabout 1 month agoApply →<h3>The Team</h3> <p>This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design &amp; build complex systems, operate with autonomy and act as owner for everything you do.&nbsp;</p> <p>The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet.&nbsp;</p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads.&nbsp;</p> <h3>Role Overview</h3> <p>We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers.&nbsp;</p> <h3>The ideal candidate should</h3> <ul> <li>Have 5+ years of experience running critical systems at scale</li> <li>Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”)</li> <li>Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment</li> <li>A strong understanding of how to run a large scale Linux environment, including low level fundamentals&nbsp;</li> <li>Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python)</li> <li>Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc)</li> </ul> <h3>Special Requirements:</h3> <ul> <li>Be a US Citizen</li> </ul> <h3>Expectations</h3> <ul> <li>Participate in the development of a reliable and resilient multi-cloud platform that hosts business critical applications for a wide &amp; varied range of customer applications</li> <li>Collaborate with service-owning teams to provide internal support, solve technical challenges and adapt or build tooling to solve novel use cases in a generic fashion</li> <li>Participate in a 24/7 on-call rotation to swiftly resolve issues related to any disruption of our customer facing Atlas fleet, ensuring minimal disruption and high availability</li> </ul> <h3>About MongoDB</h3> <p>MongoDB is built for change, empowering our customers
Site Reliability Engineer (Senior or Staff), Infrastructure Security
MongoDB · Austin; New York City; San Francisco; Seattle; United States, USA
On-siteabout 1 month agoApply →<p>We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.</p> <p>The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture.</p> <p>This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation.</p> <p>This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones.</p> <h3>Responsibilities:</h3> <p>Cloud Security Design and Implementation:&nbsp;&nbsp;</p> <ul> <li>Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM)</li> </ul> <p>Automation and Monitoring:&nbsp;&nbsp;</p> <ul> <li>Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection</li> </ul> <p>Security Tooling:</p> <ul> <li>Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM&nbsp;</li> </ul> <h3>Qualifications:</h3> <p>Experience:&nbsp;&nbsp;</p> <ul> <li>6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role</li> </ul> <p>Security Mindset:</p> <ul> <li>A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking layers to cloud provider configurations. Proven experience in leading projects within security-focused areas, such as runtime scanning, security observability, CSPM, and more</li> </ul> <p>Cloud Expertise:&nbsp;&nbsp;</p> <ul> <li>Strong experience with at least one cloud platform (AWS, Azure, GCP), including expertise in IAM, VPC networking, security groups, and cloud security tools (e.g., GuardDuty, Security Hub, CloudTrail)</li> </ul> <p>Coding/Automation:&nbsp;&nbsp;</p> <ul>
Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
MongoDB · Boston; Miami; New York City; Pittsburgh; Raleigh; United States, USA
On-siteabout 1 month agoApply →<p>MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently.</p> <p>You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture.</p> <p>This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location.&nbsp;</p> <h3>The ideal candidate should</h3> <ul> <li>Have 6+ years of experience working on software development and operating distributed systems</li> <li>Proficiency in Python, Go, or a similar language</li> <li>Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs.</li> <li>Possess a customer-focused mindset</li> <li>Value efficiency in processes and operations</li> <li>Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil</li> <li>Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market</li> <li>Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure</li> <li>Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)</li> </ul> <h3>Responsibilities&nbsp;</h3> <ul> <li>Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs</li> <li>Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing</li> <li>Identify and configure key metrics to detect incidents and quantify service health, availability, and performance</li> <li>Participate in a 24/7 on-call rotation to resolve issues involving the storage infrastructure</li> <li>Become an expert in infrastructure performance, helping us optimize from the application level all the way to the kernel</li> </ul> <h3>Strong candidates may also have experience with:</h3> <
Senior Site Reliability Engineer, Fleet Management
MongoDB · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, USA
On-siteabout 1 month agoApply →<h3>The Team</h3> <p>Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems.</p> <p>The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model.</p> <p>This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin.</p> <h3>Responsibilities</h3> <ul> <li>Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB</li> <li>Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems</li> <li>Participate in a 24/7 on-call rotation to resolve critical issues</li> <ul> <li>Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice</li> </ul> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 6+ years of experience in software development and operating distributed systems</li> <li>Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests)</li> <li>Have deep experience using and extending containerization technologies, preferably Kubernetes</li> <li>Have a solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS)</li> <li>Possess a customer focused mindset, treating internal developers as your primary users</li> <li>Have strong operational ownership, including a track record of debugging complex production issues and driving them to resolution</li> <li>Prefer automation over manual processes ("allergic to ops work")</li> <li>We are a small team of software engineers with a strong bias toward building software solutions to eliminate toil</li> </ul> <h3>Strong candidates may also have experience w
Manager, Site Reliability Engineering - Storage Layer Service
MongoDB · New York City, USA
On-siteabout 1 month agoApply →<p>MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently.</p> <p>As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture.</p> <p>We are looking to speak to candidates who are based in New York City for our hybrid working model.</p> <h3>Responsibilities</h3> <ul> <li>Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers</li> <li>Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs</li> <li>Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges</li> <li>Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations</li> </ul> <h3>You may be a good fit if you</h3> <ul> <li>Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams</li> <li>Possess a customer-focused mindset, treating internal developers as your primary users</li> <li>Value efficiency in processes and operations, and have a track record of optimizing team workflows</li> <li>Prefer automation over manual processes, fostering a culture of building software solutions to eliminate toil</li> <li>Have deep technical familiarity with Kubernetes ecosystems, containerization technologies, and modern IaC tooling (e.g., Terraform, Crossplane, or Operators) so you can effectively guide the team's technical decisions</li> <li>Have operated or supported stateful storage or database systems at scale and are comfortable with durability, consistency and recovery trade-offs</li> <li>Excel at translating complex business and engineering requirements into actionable, phased technical roadmaps</li> <li>Have a high level of empathy, res
Senior Site Reliability Engineer (In-Office Required)
nebius · New York City, United States
On-siteabout 1 month agoApply →Staff Site Reliability Engineer
Diligent Corporation · New York, USA
On-siteabout 1 month agoApply →<p><strong>Position Overview</strong></p> <p><span data-contrast="auto">We are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation frameworks, to join our global Infrastructure &amp; Operations team. This role is a hands-on senior engineering position responsible for designing, maintaining, and optimizing our private cloud environments, which underpin mission-critical SaaS products.</span><span data-ccp-props="{&quot;134233117&quot;:true,&quot;134233118&quot;:true,&quot;201341983&quot;:0,&quot;335559740&quot;:240}">&nbsp;</span></p> <p><span data-contrast="auto">The ideal candidate will have extensive experience&nbsp;operating&nbsp;in enterprise&nbsp;datacenter&nbsp;environments,&nbsp;a strong foundation&nbsp;in Microsoft Active Directory and Windows Server, and a proven ability to build—not just run—automation workflows that improve reliability, scalability, and efficiency.</span><span data-ccp-props="{&quot;134233117&quot;:true,&quot;134233118&quot;:true,&quot;201341983&quot;:0,&quot;335559740&quot;:240}">&nbsp;</span></p> <p><span data-contrast="auto">You will work closely with other engineering teams (Network, Security, SRE, and DevOps) to ensure the stability and performance of our global platform and drive continuous improvement through automation and infrastructure modernization.</span><span data-ccp-props="{&quot;134233117&quot;:true,&quot;134233118&quot;:true,&quot;201341983&quot;:0,&quot;335559740&quot;:240}">&nbsp;</span></p> <p><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:0,&quot;335559740&quot;:240}">&nbsp;</span></p> <p><strong><span data-contrast="auto">Key Responsibilities</span></strong></p> <ul> <li data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="1" data-aria-level="1"><span data-contrast="auto">Architect, deploy, and&nbsp;maintain</span><span data-contrast="auto">&nbsp;VMware-based private cloud infrastructure across multiple global&nbsp;datacenters.</span><span data-ccp-p
Senior Site Reliability Engineer (In-Office Required)
Nebius · New York City, USA
On-siteabout 1 month agoApply →<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><h3><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">About&nbsp;</span><span data-ccp-parastyle="heading 2">Tavily</span></span></strong><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335551550&quot;:0,&quot;335551620&quot;:0,&quot;335559738&quot;:160,&quot;335559739&quot;:80}">&nbsp;</span></h3> <p><span data-contrast="none">We're&nbsp;building the infrastructure layer for agentic web interaction at scale. Our API is designed from the ground up to power Retrieval-Augmented Generation (RAG) and real-time reasoning in AI systems. By connecting LLMs to high-quality, trustworthy web content, we help developers build agents that are not only intelligent — but also informed.</span><span data-ccp-props="{&quot;335551550&quot;:0,&quot;335551620&quot;:0}">&nbsp;</span></p> <p><span data-contrast="none">We work with some of the most innovative teams in AI — from small startups shaping the ecosystem to the largest enterprises deploying AI at scale. Whether&nbsp;it's&nbsp;powering sales assistants, research copilots, or internal knowledge tools,&nbsp;we're&nbsp;the missing link between LLMs and the real world.</span><span data-ccp-props="{&quot;335551550&quot;:0,&quot;335551620&quot;:0}">&nbsp;</span></p> <h3><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">The Role:</span></span></strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">&nbsp;</span><span data-ccp-parastyle="heading 2">Senior&nbsp;</span><span data-ccp-parastyle="heading 2">Site Reliability</sp
Site Reliability Engineer
Kalshi · New York Office, USA
On-siteabout 1 month agoApply →Site Reliability Engineer at Kalshi. Apply via Ashby.
Site Reliability Engineer (Senior or Staff), Deployments
mongodb · Boston; Miami; New Jersey; New York City; Princeton; Raleigh; Washington DC, Boston; Miami; New Jersey; New York City; Princeton; Raleigh; Washington DC
On-siteabout 1 month agoApply →Staff Site Reliability Engineer
diligentcorporation · New York, United States
On-siteabout 1 month agoApply →Staff Site Reliability Engineer
Legora · New York City, USA
On-siteabout 1 month agoApply →Staff Site Reliability Engineer at Legora. Apply via Ashby.
Senior Site Reliability Engineer
Legora · New York City, USA
On-siteabout 1 month agoApply →Senior Site Reliability Engineer at Legora. Apply via Ashby.