Jobs and Careers
SA

Software Engineer LMTS (Site Reliability Engineering)

Salesforce
San Francisco, United Statesfull_timeVerifiedPosted 12 Jun 2025
💰 $276,100/yr($184,000/yr$276,100/yr)

About the role

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.

Job Category

Software Engineering

Job Details

About Salesforce

We’re Salesforce, the Customer Company, inspiring the future of business with AI+ Data +CRM. Leading with our core values, we help companies across every industry blaze new trails and connect with customers in a whole new way. And, we empower you to be a Trailblazer, too — driving your performance and career growth, charting new paths, and improving the state of the world. If you believe in business as the greatest platform for change and in companies doing well and doing good – you’ve come to the right place.

This candidate must be a U.S. citizen (U.S. born or naturalized) operating on U.S. Soil who does not hold dual citizenship with the ability to meet customer and government screening standards applicable to this role.

This position requires onsite presence in either Boston, San Francisco or Bellevue offices.

As a Software Engineer in Site Reliability Engineering (SRE) at MuleSoft, you will be part of a high-impact team focused on architecting, building, and scaling the infrastructure, tools, and platforms that improve the resiliency, reliability, performance, and scalability of distributed systems running on the MuleSoft Anypoint Platform. This is a software engineering-driven role, where you'll write production-grade code to automate operations, enhance observability, and strengthen service resilience—especially in high-security environments, including FedRAMP, Protected B, among others.


Your work will span the entire stack: from shaping engineering practices and building proactive failure-prevention mechanisms to streamlining deployment pipelines and improving the end-to-end reliability of mission-critical services. As stewards of observability, incident management, release automation, and reliability engineering, our team’s mission is to embed resiliency and reliability into every layer of the system and consistently exceed industry standards for uptime, latency, and performance.

What You’ll Be Doing

  • Engineering Resiliency and Reliability: Design and develop systems, libraries, and tools that strengthen the resiliency and reliability of distributed services running on the MuleSoft Anypoint Platform.
  • Observability by Design: Develop and extend monitoring, logging, and alerting capabilities using industry-standard observability platforms (e.g., metrics, tracing, and log aggregation tools) to ensure issues are detected and diagnosed before they impact customers.
  • Automation at Scale: Write production-grade code in Python, Go, or similar languages to automate operational tasks, scale deployment pipelines, and implement self-healing systems.
  • Incident Response & Prevention: Participate in on-call rotations, drive root cause analysis, and deliver software-based solutions that prevent recurrence and reduce meantime to recovery (MTTR).
  • Platform and Infrastructure Development: Build internal platforms, shared APIs, and systems that enhance developer velocity while improving overall system resilience and operability.
  • CI/CD and Deployment Engineering: Optimize and evolve our CI/CD pipelines using Jenkins, Spinnaker, and infrastructure-as-code tools such as Terraform and Kubernetes to enable safe and frequent delivery.
  • Security and Compliance as Code: Develop and maintain automated solutions to meet FedRAMP, Protected B, and other regulatory requirements—integrating security and compliance directly into deployment workflows.
  • Collaborative Reliability Advocacy: Work closely with product engineers, platform teams, and security stakeholders to influence architectural decisions and bake reliability into all layers of the stack.
  • Runbooks and Design Documentation: Create and maintain high-quality documentation for systems, processes, and playbooks to promote operational excellence and team scalability.


Requirements:

  • 8+ years of experience in Software Engineering, SRE, or DevOps roles, with a strong focus on building resilient, scalable, and highly available systems.
  • Proven proficiency in Java, Python, Go, Bash, with experience writing production-quality, maintainable, and testable code for infrastructure and platform automation.
  • Hands-on experience with infrastructure as code, CI/CD pipelines, and deployment automation using tools like Terraform, Jenkins, and Spinnaker.
  • Proven experience architecting, developing, and operating systems in cloud-native environments (AWS) and managing containerized workloads with Kubernetes.
  • Strong unde

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Salesforce

View company profile →