Jobs and Careers
SA

Director, Site Reliability Engineering

Salesforce
New York City, United Statesfull_timeVerifiedPosted 4 Aug 2026
💰 $313,700/yr($197,300/yr$313,700/yr)

About the role

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.

Job Category

Software Engineering

Job Details

About Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.

Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

Job Title: Director, Site Reliability Engineering
Location: New York, NY; San Francisco, CA; Dallas, TX

About the Role

We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.

In this role, you will transform our SRE function—moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.

As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.

Key Responsibilities

SRE Strategy and Leadership

  • Define and execute the long-term strategy and roadmap for Site Reliability Engineering.

  • Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.

  • Build and develop a high-performing team of site reliability and operations engineers.

  • Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.

  • Translate business priorities and customer impact into clear reliability investments and engineering outcomes.

  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.

Reliability Engineering

  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.

  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.

  • Define what it means for a service to be operationally and observably ready for production.

  • Develop readiness reviews and certification practices for high-impact services and launches.

  • Drive improvements in system availability, performance, resiliency, and recovery.

  • Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.

Observability

  • Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.

  • Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.

  • Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.

  • Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.

  • Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.

  • Establish governance and measurement to assess adoption and effectiveness of observability standards.

Incident Management and Operational Excellence

  • Improve incident detection, response, mitigat

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Salesforce

View company profile →