Staff-Software Engineer
American ExpressAbout the role
At American Express, we are transforming how software is built, deployed, and operated at enterprise scale. The Site Reliability Engineering (SRE) organization is at the center of this transformation, building the engineering capabilities that enable secure software delivery, resilient platforms, intelligent operations, and exceptional developer experiences.
As a Staff Engineer, you will be a senior technical leader responsible for advancing the reliability, scalability, and operational excellence of critical technology platforms. You will work across engineering organizations to solve complex production challenges through software engineering, platform innovation, automation, and AI-powered operational capabilities.
This is a highly influential individual contributor role where success is measured by the systems you build, the engineering practices you shape, and the enterprise impact you create.
Role Summary
As a Staff Engineer within the Site Reliability Engineering organization, you will lead the evolution of enterprise reliability engineering by building scalable platform capabilities that improve system resilience, engineering productivity, and operational efficiency.
You will partner closely with Platform Engineering, Infrastructure Engineering, Security Engineering, Application Development, and Enterprise Architecture teams to modernize how software is delivered, observed, secured, and operated across the enterprise.
This role requires deep expertise in distributed systems, cloud infrastructure, software engineering, DevSecOps, observability, automation, and AI-enabled operations. You will drive engineering solutions that reduce operational complexity, improve service reliability, eliminate manual toil, and accelerate the organization's journey toward autonomous operations.
Key Responsibilities
Lead Enterprise Reliability Engineering
Drive the technical direction for reliability engineering across critical enterprise platforms.
You will:
- Define and evolve enterprise-wide Site Reliability Engineering practices and standards.
- Champion Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to improve service reliability and engineering accountability.
- Establish engineering patterns for high availability, scalability, resiliency, and disaster recovery.
- Lead architecture and design reviews for mission-critical platforms and distributed systems.
- Partner with engineering teams to improve production readiness, operational maturity, and service resilience.
- Drive adoption of proactive reliability engineering practices, including capacity planning, performance optimization, failure testing, and resilience validation.
- Build Software That Improves Operations
Apply software engineering to eliminate operational complexity and improve engineering productivity.
You will:
- Design and develop reusable platforms, frameworks, and automation that reduce operational toil.
- Build engineering capabilities that simplify production operations and enable self-service experiences.
- Develop reference implementations and engineering libraries adopted across multiple organizations.
- Improve deployment safety through automated validation, policy enforcement, and release controls.
Leverage modern programming languages and cloud-native technologies to solve complex operational problems at scale.
Modernize Software Delivery & Engineering Platforms
Enable secure, reliable, and efficient software delivery across the enterprise.
You will:
- Drive maturity of enterprise CI/CD platforms through standardized pipelines and reusable engineering capabilities.
- Embed security, reliability, testing, and compliance directly into software delivery workflows.
- Advance Platform Engineering and Internal Developer Platforms (IDPs) that improve developer experience and engineering velocity.
- Improve deployment reliability using progressive delivery, automated verification, rollback strategies, and policy-as-code.
Partner with application teams to simplify software delivery while improving production stability.
Advance Intelligent Operations & Autonomous Engineering
Lead the evolution of modern operations through AI-driven engineering and intelligent automation.
You will:
- Build next-generation operational capabilities using AI, machine learning, and intelligent automation.
- Integrate Generative AI and agentic AI into incident management, diagnostics, troubleshooting, and engineering workflows.
- Design self-healing operational capabilities that automate detection, analysis, and remediation of produc
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s