Technology, DevOps/Site Reliability Engineer
BTIG · New York, USA
On-siteabout 1 hour agoApply →<p><strong>Job Purpose:</strong></p> <p>BTIG seeks a DevOps/Site Reliability Engineer to join our technology team. This role is central to improving developer velocity by handling production application support escalations, managing and evolving our infrastructure stack, and providing operational continuity across the team. The ideal candidate thrives at the intersection of software operations, infrastructure engineering, and customer-facing support and is energized by the opportunity to progressively take ownership of critical platform systems within a financial services environment.<br>In this role, you will be a force multiplier for a focused engineering team, the person who keeps production running smoothly, evolves the platform, and creates the space for developers to build. You'll gain broad, hands-on exposure across the full stack in an environment where reliability directly impacts trading operations.</p> <p><strong>Duties &amp; Responsibilities:</strong></p> <p>• Serve as the primary point of contact for production application escalations — triage, diagnosis, and resolution across application components and services.<br>• Monitor application and infrastructure health; investigate anomalies and remediate without pulling developers off feature work.<br>• Develop and maintain runbooks, escalation procedures, and operational knowledge base documentation.<br>• Identify recurring issues and collaborate with developers to drive root-cause fixes<br>• Own incident response and post-incident review processes<br>• Manage, maintain, and improve the team's infrastructure stack:<br>o Reverse proxy &amp; traffic management<br>o Identity &amp; Access Management<br>o Secret/Configuration Management<br>o Certificate management<br>o SQL and No SQL Databases<br>o OTEL/Metrics, Traces, Logs &amp; Analytics<br>o Container Orchestration<br>o Event Streaming</p> <p>• Automate provisioning, deployment, and configuration management<br>• Plan and execute upgrades, patches, security hardening, and capacity management<br>• Evolve infrastructure toward greater reliability, scalability, and developer self-service<br>• Own the observability stack — maintain and improve monitoring, alerting, dashboards, and centralized logging to ensure production issues are detected quickly and diagnosed efficiently<br>• Work with developers to ensure applications emit useful metrics, structured logs, and trace context<br>• Serve as backup to the customer support role during PTO, absences, or peak demand<br>• Develop familiarity with customer workflows, common issues, and resolution paths<br>• Contribute to support documentation and self-service tooling that reduces overall support burden</p> <p><strong>Requirements &amp; Qualifications:<
Senior Site Reliability Engineer
Dualentry · New York City, USA
On-siteabout 15 hours agoApply →Senior Site Reliability Engineer at Dualentry. Apply via Ashby.
Senior Site Reliability Engineer
dualentry · New York City, USA
On-site5 days agoApply →Site Reliability Engineer
virtu · New York, New York
On-site5 days agoApply →Senior Site Reliability Engineer
symphony · New York, New York
On-site5 days agoApply →Executive Director - Site Reliability Engineering - WM Technology
Morgan Stanley · New York, NY, United States
On-site6 days agoApply →<p>Morgan Stanley is a leading global financial services firm providing a wide range of investment banking, securities, investment management and wealth management services. The Firm's employees serve clients worldwide including corporations, governments, and individuals from more than 1,200 offices in 43 countries.<br><br>As a market leader, the talent and passion of our people is critical to our success. Together, we share a common set of values rooted in integrity, excellence, and strong team ethic. Morgan Stanley can provide a superior foundation for building a professional career - a place for people to learn, to achieve and grow. A philosophy that balances personal lifestyles, perspectives and needs is an important part of our culture.<br><br>Morgan Stanley Wealth Management Technology (WMT) is the global technology department responsible for the design, development, delivery and support of the technical solutions behind the products and services used by the Morgan Stanley Wealth Management (MSWM) business. The department is comprised of 10 organizations: Sales, Banking & Corporate-Client Technology, Investment Products & Markets Technology, Client Reporting, Core Processing, Private and International Wealth Management Technology, Technology Integration Office, Enterprise Infrastructure & Production Management, Capital Markets Application & Data Services, Deployment Planning & Release Management, and the Chief Operating Office.<br><br>The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.<br><br>The Reliability Operations team partners closely with WM business teams as well as the broader WM Tech organization and firmwide end user, network and risk technology organizations. The team is responsible for maintaining the stability, resilience, and performance of the technology platforms that support WM business. This role focuses on ensuring systems deliver first class client, investment and regulatory outcomes consistently, resiliently, sustainably and effectively.<br><br>Primary Responsibilities<br><ul><li>Own the end-to-end health and stability of production applications used across the WM Tech footprint intraday and in overnight batches (Architecture, Operations, Risk)</li><li>Lead incident, problem, and crisis management including cross-functional war rooms and senior-level communication.</li><li>Drive implementation of operational controls, monitoring, observability, and capacity-management practices.</li><li>Oversee change management, readiness assessments, and production transition for new technology capabilities.</li><li>Establish and manage SLAs, KPIs, and operational dashboards to track technology performance and provide high degree of transparency to the business.</li><li>Lead root-cause analysis and ensure remediation plans are completed and measured.</li><li>Build strong partnerships with the broader IM Tech organization and the MSIM business organization.</li><li>Manage and develop a cross regional team and foster a culture of accountability and operational excellence.</li><li>Oversee vendor performance and adherence to service standards.</li></ul><br><br>Skills Required<br><ul><li>Bachelor's degree, with 12-18+ years in technology production management or related roles in financial services.</li><li>Deep understanding of front-to-back asset-management workflows, trading systems, and data platforms.</li><li>Demonstrated leadership in mission-critical technology operations and incident management.</li><li>Demonstrated leadership in driving SRE practices across a technology stack leading to improved production stability.</li><li>Demonstrated high degree of taking accountability for client outcomes, rigor, humility & partnership.</li><li>Strong understanding of global client-service expectations and regulatory frameworks.</li><li>Excellent communication skills with experience presenting to senior leadership.</li><li>Strong organizational skills and ability to manage multiple priorities.</li><li>Ability to lead cross-regional teams and influence stakeholders globally.</li><li>Familiarity with monitoring, observability, automation, and cloud/DevOps concepts preferred.</li></ul><br><br><b>WHAT YOU CAN EXPECT FROM MORGAN STANLEY:</b> <br><br>At Morgan Stanley, we raise, manage and allocate capital for our clients - helping them reach their goals. We do it in a way that's differentiated - and we've done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren't just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you'll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There's also ample opportunity to move about the business for those who show passion and grit in their work. <br><br>To learn more about our offices across the globe, please copy and paste https://www.morganstanley.com/about-us/global-offices into your browser.<br><br>Expected base pay rates for the role will be between $195,000 and $275,000 per year at the commencement of employment. However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs.<br><br>Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents. <br><br>Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences. <br><br>For more information, please visit : https://www.morganstanley.com/people-opportunities/eeo .</p>
Lead Site Reliability Engineer
JPMorgan Chase · New York, NY, United States
On-site6 days agoApply →<p>As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology, liquidity risk team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your team's strategic planning, driving continual improvement in customer experience, resiliency, security, scalability, monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.<br><br><strong>J</strong><strong>ob responsibilitie</strong>s<br><br><ul><li>Lead SRE practices that balance delivery speed, efficiency, and system stability </li><li>Partner with engineering peers and senior stakeholders to drive strong, shared outcomes </li><li>Scale SRE adoption across application and platform teams </li><li>Set reliability expectations and show progress through stability and reliability metrics </li><li>Run blameless, data-driven post-incident reviews and regular debriefs to turn lessons into improvements </li><li>Build a continuous-improvement culture by gathering feedback and improving the customer experience </li><li>Coach entry- to mid-level engineers and promote knowledge sharing through internal forums and communities</li><li><strong>Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.</strong></li><li><strong>Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.</strong></li></ul><br>Required qualifications, capabilities, and skills<br><br><ul><li>Formal training or certification in software engineering concepts plus <strong>5+ years</strong> of applied experience </li><li>Advanced knowledge of SRE principles and a track record of implementing SRE across application and platform teams while avoiding common pitfalls </li><li>Experience leading technologists to manage and resolve complex technology issues at a firmwide level </li><li>Ability to influence team culture by championing innovation and driving change </li><li>Experience hiring, developing, and recognizing talent </li><li>Proficiency in at least one programming language (preferred: <strong>JavaScript, Go, Python</strong>) </li><li>Hands-on experience with CI/CD tools (e.g., Jenkins, GitLab, Terraform) </li><li>Experience with containers and orchestration (e.g., Docker, Kubernetes, ECS)</li><li><strong>Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.</strong></li><li><strong>Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.</strong></li></ul><br>Preferred qualifications, capabilities, and skills<br><br><ul><li>Ability to code, troubleshoot, and demonstrate strong data fluency</li><li>Strong troubleshooting skills across common networking technologies and issues </li><li>Working knowledge of modern service and integration patterns, including <strong>GraphQL fundamentals</strong>, <strong>event-driven architecture (Kafka or equivalent)</strong>, and <strong>observability/telemetry with OpenTelemetry</strong></li></ul> <br><br><b>ABOUT US</b><br><br>Chase is a leading financial services firm, helping nearly half of America's households and small businesses achieve their financial goals through a broad range of financial products. Our mission is to create engaged, lifelong relationships and put our customers at the heart of everything we do. We also help small businesses, nonprofits and cities grow, delivering solutions to solve all their financial needs.<br><br>We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.<br><br>We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.<br><br>Equal Opportunity Employer/Disability/Veterans <br><br><b>ABOUT THE TEAM</b><br><br>Our Consumer & Community Banking division serves our Chase customers through a range of financial services, including personal banking, credit cards, mortgages, auto financing, investment advice, small business loans and payment processing. We're proud to lead the U.S. in credit card sales and deposit growth and have the most-used digital solutions - all while ranking first in customer satisfaction.</p>
Site Reliability Engineer, Pragma
marketaxesscorporation · New York, United States
On-site12 days agoApply →Lead Site Reliability Engineer
movableink · New York, United States
On-site14 days agoApply →Staff Site Reliability Engineer
tabs · New York City, NY, USA
On-site16 days agoApply →Site Reliability Engineer
claylabs · New York, USA
On-site16 days agoApply →Director, Site Reliability Engineering
stellar · New York, USA
On-site16 days agoApply →Staff Site Reliability Engineer
legora · New York City, USA
On-site16 days agoApply →Senior Site Reliability Engineer
legora · New York City, USA
On-site16 days agoApply →Site Reliability Engineer
kalshi · New York Office, USA
On-site17 days agoApply →Senior Software Engineer, Site Reliability Engineering
Ridgeline · New York, USA
On-site19 days agoApply →<h3 class="PDq2pG_selectionAnchorContainer" data-start="556" data-end="1063">Senior Software Engineer, Site Reliability Engineering</h3> <h4 class="PDq2pG_selectionAnchorContainer" data-start="556" data-end="1063">Reno, NV; San Ramon, CA; NYC - Hybrid</h4> <p class="PDq2pG_selectionAnchorContainer" data-start="556" data-end="1063">Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating complex operational challenges, improving observability, and eliminating manual toil through thoughtful engineering? Are you excited by the opportunity to support mission-critical production systems while collaborating with talented engineers in a fast-moving, innovative environment? If so, we invite you to be a part of our innovative team.</p> <p data-start="1065" data-end="1789">As a Site Reliability Engineer, you'll help ensure the reliability, scalability, and operational excellence of Ridgeline's mission-critical SaaS platform. You'll partner closely with product and platform engineers to improve service reliability, accelerate engineering velocity through automation, and build systems that are easier to operate from day one. Our team of engineers are building with cutting-edge technologies—like Claude Code and Cursor—in a fast-moving, creative, progressive work environment. You'll play a key role in advancing our observability, release engineering, incident response, and automation capabilities while contributing measurable improvements to platform stability and developer productivity.</p> <p data-start="1791" data-end="2179">At Ridgeline, how we work matters as much as what we build. Ridgeliners act like owners, choose growth over comfort, and communicate with transparency. We assume positive intent, bias toward action, and bring solutions—not just problems. We celebrate wins, learn from setbacks, and thrive in a resilient, collaborative, high-performing culture. If this excites you, we'd love to meet you!</p> <p data-start="2181" data-end="2276"><strong data-start="2181" data-end="2276">You must be work authorized in the United States without the need for employer sponsorship.</strong></p> <h2 data-section-id="12o973t" data-start="2278" data-end="2309"><strong data-start="2281" data-end="2309">The impact you will have</strong></h2> <ul data-start="2311" data-end="3637"> <li data-section-id="1kzrxan" data-start="2311" data-end="2425">Improve the reliability, availability, and performance of Ridgeline's mission-critical production SaaS platform.&
Senior Site Reliability Engineer, Axon 911
axon · New York, United States
On-site19 days agoApply →Senior Site Reliability Engineer, Observability
ripple · New York, United States
On-site20 days agoApply →Senior / Staff Site Reliability Engineer
radar · New York , USA
On-site21 days agoApply →Technical Product Manager II, Site Reliability Engineering
The New York Times · New York, USA
On-site21 days agoApply →<div class="content-intro"><div id="labeledImage.LOCATION" class="WIFG" data-automation-id="decorationWrapper"> <div class="WNHJ"> <div id="labeledImage.LOCATION--uid38" class="WE-Y WMXY WBAB WF0Y" data-automation-id="responsiveMonikerInput" data-metadata-id="labeledImage.LOCATION" data-uxi-form-item-child-list-index="0"> <div class="WJ-Y"> <p><strong>The <a href="https://www.nytco.com/company/mission-and-values/" target="_blank"><u>mission</u></a> of The New York Times is to seek the truth and help people understand the world. That means independent journalism is at the heart of all we do as a company. It’s why we have a world-renowned newsroom that sends journalists to report on the ground from nearly 160 countries. It’s why we focus deeply on how our readers will experience our journalism, from print to audio to a world-class digital and app destination. And it’s why our business strategy centers on making journalism so good that it’s worth paying for.&nbsp;</strong></p> </div> </div> </div> </div></div><p><strong>Mission Overview &amp; Responsibilities</strong><span style="font-weight: 400;">:&nbsp;</span></p> <p data-pm-slice="1 1 []">At The New York Times, our Site Reliability Engineering (SRE) team is central to how we design, test, and operate the systems that support our most critical customer experiences. We're looking for a Technical Product Manager to lead the strategy for reliability programs and platforms that help teams ship resilient systems with confidence. These programs include operational readiness, load and chaos testing, observability, and incident readiness.</p> <p>You'll partner with SRE, platform infrastructure, and product engineering teams to define the standards, tooling, and practices that improve operational readiness across hundreds of services. You will focus on building scalable reliability programs and experiences—not running cloud infrastructure or managing operational tickets.</p> <p>You will be the product lead for a portfolio of SRE programs that includes:</p> <ul> <li> <p>Operational Readiness and Always Ready – reliability models, scorecards, production readiness reviews, and reliability signals that support operating reviews and critical customer journeys.</p> </li> <li> <p>Load, Chaos &amp; Disaster Recovery Testing – platforms and practices that validate how systems behave under high traffic, zonal or regional failures, and degraded conditions.</p> </li> <li> <p>Observability &amp; Incident Readiness – opinionated defaults for metrics, logs, traces, dashboards, alerts, and runbooks that i