Site Reliability Engineer Jobs in New York, USA

72 verified site reliability engineer openings in New York

  • Director of Site Reliability Engineering

    Stellar · New York, USA

    On-site
    about 1 month agoApply →

    Director of Site Reliability Engineering at Stellar. Apply via Ashby.

  • Senior Site Reliability Engineer, Data Infrastructure

    CoreWeave · New York, USA

    On-site
    about 1 month agoApply →

    <div class="content-intro"><div> <div> <div class="gmail_quote"> <div> <div><span id="m_1770241969069985273m_-2746164444908759431gmail-docs-internal-guid-131e4fb0-7fff-b4e9-ff50-e8cf32449b1b">CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at <a href="http://www.coreweave.com/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=http://www.coreweave.com&source=gmail&ust=1762613132717000&usg=AOvVaw3D-UOhNaqEvF5BEWxjYyAU">www.coreweave.com</a>.</span></div> </div> </div> </div> </div></div><p><strong>What You’ll Do:</strong></p> <p>The Platform & Infrastructure Engineering team in the Data Infrastructure organization is responsible for the reliability, scalability, and security of the company’s data platform. The team builds and operates the foundational systems that power data ingestion, transformation, analytics, and internal AI workloads at scale. We operate with production-grade discipline, supporting mission-critical services with stringent uptime requirements and a focus on automation, observability, and resilience.</p> <p><strong>About the role:</strong></p> <p>As a Senior Site Reliability Engineer, you will own the reliability and performance of our Kubernetes-based data platform. You will design and operate highly available, multi-region systems, ensuring our services meet strict uptime and latency targets. Day-to-day, you’ll work on scaling infrastructure, improving deployment pipelines, and hardening our security posture. You’ll play a key role in evolving our DevSecOps practices while partnering closely with engineering teams to ensure services are built for reliability from day one.</p> <p><strong>Who You Are: </strong></p> <ul> <li>5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering roles</li> <li>Deep expertise in Kubernetes and containerized software services, including cluster design, operations, and troubleshooting in production environments</li> <li>Strong experience building and operating CI/CD systems, including tools such as Argo CD and GitHub Actions</li> <li>Proven experience owning production systems with high availability requirements (≥99.99% uptime), inclu

  • Site Reliability Engineer

    Peloton · New York, USA

    On-site
    about 2 months agoApply →

    <p><strong>ABOUT THE ROLE</strong></p> <p>At Peloton, we view Platform as a Product. A phenomenal platform unlocks speed of development and learning. It allows us to scale easily, enabling our engineers to maximize attention on new features and capabilities.  A key to crafting a phenomenal platform is data-driven insights and understanding where we should focus our attention to create the best outcomes for our members.  Platform at Peloton is a force-multiplier that enables Peloton to move faster and scale safely with minimal effort. Core to this mission is creation of the best developer experience in the tech industry for the entire spectrum of Peloton's technology.  We work across an incredible range of technology domains: hardware, firmware, web, mobile, backend, data, messaging, content, streaming, and machine learning. We get to apply these to create a platform of products loved by millions of customers all over the world. </p> <p>Peloton is looking for a Site Reliability Engineer with an operations focus to work with teams across the organization to help build and maintain a monitorable, performant, reliable, and highly-scalable deployment platform.  We are a growing team of engineers tackling exciting problems to handle thousands of nodes and pods spread across many deployments.</p> <p><strong>YOUR DAILY IMPACT AT PELOTON</strong></p> <ul> <li>Automatic, fast auto scaling for live rides and special large events</li> <li>Host a critical infrastructure that ensures that our members have the best experience possible on tens of thousands of pods across multiple clusters</li> <li>Provide a platform for machine learning (and other awesome workloads)  Allow developers to move quickly and experiment, without getting in the way</li> <li>Promote best practices for building and operating highly reliable systems</li> <li>Serve as domain expert in observability and monitoring</li> <li>Consult in system design to meet reliability and capacity requirements</li> <li>Automate everything, from infrastructure down to day-to-day tasks</li> <li>Conduct timely post-mortems of infrastructure incidents</li> <li>Assist with all aspects of operational security and compliance</li> <li>Seek out potential threats to security and reliability and advocate solutions</li> <li>We work with Amazon Web Services, Chef, Python, Ubuntu, Nginx, Jenkins, and Terraform</li> </ul> <p> <strong>YOU BRING TO PELOTON</strong></p> <ul> <li>Experience maintaining scalable and stable Kubernetes clusters</li> <li>Knowledge of best practices when it comes to the observability and monitoring required of running Kubernetes at<br>scale</li> <li>Knowle

  • Staff Site Reliability Engineer

    Ro · New York, NY

    On-site
    about 2 months agoApply →

    Join Tech @ Ro to build the future of healthcare, from the ground up! At Ro, we believe that when people achieve their health goals, they can achieve their life goals. The highest-leverage way to move society forward is to give people their health, and the current healthcare system isn’t built to do that. It was built to bill, not to serve patients. We’re building a new system. One where the patient is in control. One designed from scratch for the digital age. At Ro, technology isn’t just a function… It's core to how we deliver care. We’ve built a vertically integrated healthcare platform that connects telehealth, diagnostics, pharmacy, and logistics into a seamless, end-to-end experience for millions of patients. …and we’re just getting started.  As part of Tech @ Ro, you’ll work on systems that operate at scale, with an opportunity to: Solve complex, high-concurrency problems across a full-stack platform Build and ship quickly with tight feedback loops and real-world impact Own systems end-to-end, from architecture to production performance Work alongside experienced operators, technical leaders, and clinicians Help define how modern healthcare should be delivered We’re a performance-driven team with a strong sense of ownership and urgency. We move fast, learn quickly, and hold a high bar for what we build, and do so with a big heart — because patients depend on it. If you’re motivated by impact, scale, and the chance to help lead the patient revolution, come build with us. The Role:   At Ro, our mission is to provide world-class healthcare by putting patients first - and that mission depends on reliable, secure, and scalable systems. As a Staff SRE on the infrastructure team, you’ll sit at the core of that effort: owning the reliability of our production systems, hardening infrastructure and building tools that empower our engineers to ship safely and confidently.   You will work across teams to drive uptime, performance and observability – partnering closely with product, platform and security engineers.    From designing resilient systems to shaping incident response practices, this is a role for engineers who thrive on impact and care deeply about operational excellence.

  • Senior Site Reliability Engineer

    Ro · New York, NY

    On-site
    about 2 months agoApply →

    Join Tech @ Ro to build the future of healthcare, from the ground up! At Ro, we believe that when people achieve their health goals, they can achieve their life goals. The highest-leverage way to move society forward is to give people their health, and the current healthcare system isn’t built to do that. It was built to bill, not to serve patients. We’re building a new system. One where the patient is in control. One designed from scratch for the digital age. At Ro, technology isn’t just a function… It's core to how we deliver care. We’ve built a vertically integrated healthcare platform that connects telehealth, diagnostics, pharmacy, and logistics into a seamless, end-to-end experience for millions of patients. …and we’re just getting started.  As part of Tech @ Ro, you’ll work on systems that operate at scale, with an opportunity to: Solve complex, high-concurrency problems across a full-stack platform Build and ship quickly with tight feedback loops and real-world impact Own systems end-to-end, from architecture to production performance Work alongside experienced operators, technical leaders, and clinicians Help define how modern healthcare should be delivered We’re a performance-driven team with a strong sense of ownership and urgency. We move fast, learn quickly, and hold a high bar for what we build, and do so with a big heart — because patients depend on it. If you’re motivated by impact, scale, and the chance to help lead the patient revolution, come build with us. The Role At Ro, our mission is to provide world-class healthcare by putting patients first - and that mission depends on reliable, secure, and scalable systems. As a Senior SRE on the infrastructure team, you’ll sit at the core of that effort: contributing to the reliability of our production systems, hardening infrastructure and building tools that empower our engineers to ship safely and confidently.   You will work across teams to drive uptime, performance and observability – partnering closely with product, platform and security engineers. From designing resilient systems to shaping incident response practices, this is a role for engineers who thrive on impact and care deeply about operational excellence.

  • Site Reliability Engineer - NYC

    Mistral · New York, New York

    On-site
    about 2 months agoApply →

    About Mistral    At Mistral AI, we believe in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to integrate seamlessly into daily working life.   We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet enterprise needs, whether on-premises or in cloud environments. Our offerings include le Chat, the AI assistant for life and work.   We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.   Join us to be part of a pioneering company shaping the future of AI. Together, we can make a meaningful impact. See more about our culture on https://mistral.ai/careers.   Role Summary    We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and exceed our internal and external customers' expectations.     What you will do   As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.   Operations • Design, build, and maintain scalable, highly available and fault-tolerant infrastructures to support our web services and ML workloads • Make sure our platform, inference and model training environments are always highly available and enable seamless replication of work environments across several HPC clusters • Operate systems and troubleshoot issues in production environments (interrupts, on-call responses, users admin, data extraction, infrastructure scaling, etc.) • Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our client-facing APIs and large training runs • Participate occasionally in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences   Development • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform • Collaborate with AI/ML researchers to develop and implement solutions that enable safe and reproducible model-training experiments • Build a cloud-agnostic platform offering an abstraction layer between science and infrastructure • Design and develop new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.) • Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements • Document processes and procedures to ensure consistency and knowledge sharing across the team • Contribute to open-source projects, research publications, blog articles and conferences   About you   • Master’s degree in Computer Science, Engineering or a related field • 7+ years of experience in a DevOps/SRE role • Strong experience with cloud computing and highly available distributed systems • Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...) • Experience working against reliability KPIs (observability, alerting, SLAs) • Hands-on experience with CI/CD, containerization and orchestration tools (Docker, Kubernetes...) • Knowledge of monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack, Datadog...) • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices • Strong understanding of networking, security, and system administration concepts • Excellent problem-solving and communication skills • Self-motivated and able to work well in a fast-paced startup environment   Your application will be all the more interesting if you also have: • experience in an AI/ML environment • experience of high-performance computing (HPC) systems and workload managers (Slurm) • worked with modern AI-oriented solutions (Fluidstack, Coreweave, Vast...)   By applying, you agree to our Applicant Privacy Policy.  

  • Senior Site Reliability Engineer

    Asapp 2 · New York, New York

    On-site
    about 2 months agoApply →

    At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we’re guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed, ownership, and a relentless focus on outcomes. We work in tight, skilled teams, prioritize clarity over complexity, and continuously evolve through curiosity, data, and craftsmanship. We’re seeking technologists and problem solvers who thrive in fast-paced environments, love collaborating with great talent, and approach every day like it’s Day 1.   We're a globally diverse team with hubs in New York City, Mountain View, Latin America, and India—embracing both hybrid and remote work to bring the best minds together, wherever they are. If you're driven by continuous learning, rapid pivots, and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns the entire infrastructure stacks. SREs design and implement the tools that automate building reliable and performant systems. We emphasize building tools over manual processes. We implement, not administer.  We’re obsessed with automation, not repetition. Our job is to focus on building reliable infrastructure and tools for our product teams so that they can solve customer problems and deliver new features, not reinvent platforms.

  • Site Reliability Engineer

    Claylabs · New York, USA

    On-site
    about 2 months agoApply →

    Site Reliability Engineer at Claylabs. Apply via Ashby.

  • Site Reliability Engineer

    Pico · New York City, USA

    On-site
    about 2 months agoApply →

    <div class="content-intro"><p>Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions.  As financial industry experts at the center of markets and technology, we help our clients efficiently scale their business and quickly access markets. From infrastructure to connectivity, we support our clients through the full trading lifecycle.  We are a global company headquartered in New York, with offices in Chicago, London, Singapore, Hong Kong and Tokyo.</p> <p> </p></div><p><strong>Purpose of the role</strong></p> <p>Pico Site Reliability Engineering is a customer-faced group engaged with our customers, development and production teams to ensure customer success with implementing Redline products. Preferred candidates strive to work in an entrepreneurial environment, have an interest in automation and technology, and are committed to customer satisfaction.</p> <p><strong>Responsibilities and duties </strong><em>(include but not limited to)</em></p> <ul> <li>Assist clients in solving problems and maximizing the value of Redline products and services within their enterprise.</li> <li>Respond promptly to production issues, provide workarounds whenever possible.</li> <li>Replicate customer issues and verify fixes for customer patches and releases.</li> <li>Manage production environment:</li> </ul> <ul> <ul> <li>Software upgrade</li> <li>Routine changes</li> <li>Proactive monitoring</li> <li>Performance tuning of applications</li> <li>Assess impact of exchange-driven initiatives and participate in relevant weekend testing</li> </ul> </ul> <ul> <li>Work closely with development engineers to promptly resolve customer requests.</li> <li>Work closely with Pico server and network teams for issue troubleshooting and change management.</li> <li>Elevate Redline managed practice and ensure customer success by:</li> </ul> <ul> <ul> <li>Putting in place the necessary automation and scripting.</li> <li>Continuous improvement of tools (monitoring, inventory, runbooks.)</li> </ul> </ul> <p><strong>Education, Skills and background </strong>(incl. Education and Experience Requirements)</p> <ul> <li>Bachelor’s degree or higher in an Engineering discipline ie. Computer Science, Electrical Engineering, Math or Management Information Systems or relevant work experience. </li> <li>Experience with financial services front office systems, market data tickerplants, order execution technology and the global exchange technology landscape.</li&

  • Site Reliability Engineer - Enterprise Technology

    wehrtyou · New York, United States

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer, Storage - Enterprise Technology

    wehrtyou · New York, United States

    On-site
    about 2 months agoApply →
  • Technical Product Manager II, Site Reliability Engineering

    thenewyorktimes · New York, NY

    On-site
    about 2 months agoApply →
  • Senior/Staff Site Reliability Engineer

    sage49 · New York, United States

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer, Observability

    ripple · New York, United States

    On-site
    about 2 months agoApply →
  • Site Reliability Engineering Manager

    rapidsos · Boston or New York, Boston or New York

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer

    picoquantitativetrading · New York City, NY

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer

    peloton · New York, New York

    On-site
    about 2 months agoApply →
  • Manager, Site Reliability Engineering - Fleet Management

    mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle

    On-site
    about 2 months agoApply →
  • Senior Site Reliability Engineer, Fleet Management

    mongodb · Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States, Austin; Boston; Chicago; Denver; Miami; New York City; San Francisco; Seattle; United States

    On-site
    about 2 months agoApply →
  • Site Reliability Engineer 3

    mongodb · New York City, New York City

    On-site
    about 2 months agoApply →