Site Reliability Engineer, Global Banking & Markets, Vice President
Goldman SachsAbout the role
What We Do
At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets.
Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience, and performance of core business services that underpin a global 24×7 trading operation. Working across Global Markets' front, middle, and back office functions, you will engineer reliability while balancing stringent non-functional demands for availability, latency, and resilience as well as complex, evolving business requirements.
Want to push the limit of digital possibilities? Start here.
Who We Look For
Goldman Sachs Engineers are at the forefront of innovation, driving solutions as creative collaborators in a fast-paced global environment. We seek individuals who evolve, adapt, and thrive on challenging problems.
As part of our SRE team, you will operate at the intersection of reliability engineering, cloud infrastructure, and AI-driven operations. Using Goldman Sachs' AI tooling and agentic assistants, you will accelerate incident diagnosis, automate operational toil, comprehend large legacy codebases, and raise the bar for production-quality automation across the software and reliability lifecycle. Above all, you will bring strong risk acumen and the ability to connect the right people across the organization to resolve problems quickly and decisively.
Your Impact
- Own reliability outcomes: Define and defend Service Level Objectives (SLOs), error budgets, and reliability standards for critical trading services, with risk always front of mind.
- Reduce risk and toil: Identify systemic risks before they materialize, automate away repetitive operational work, and strengthen the resilience posture of the platform.
- Connect and communicate: Act as a trusted coordinator during incidents — rapidly mobilizing the right engineers, domain experts, and stakeholders across a globally distributed organization, and communicating clearly with both technical and non-technical audiences.
- Multiply your output with AI: Orchestrate AI coding and operations agents to accelerate root-cause analysis, remediation, and automation while maintaining mastery, quality, and production fitness over all AI-generated work.
- Build for the future: Design and operate high-availability, multi-region, event-driven services on a modern cloud-native platform, setting the reliability and architectural standard for years to come.
What You Will Do
- Design, build, and operate high-availability, multi-region, cloud-native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer.
- Establish and manage SLIs, SLOs, and error budgets; drive blameless post-incident reviews and translate findings into durable engineering improvements.
- Lead incident response for latency-sensitive, high-throughput trade lifecycle systems — quickly diagnosing issues, coordinating cross-functional responders, and communicating status to stakeholders.
- Develop event-driven architectures, multi-stage processing pipelines, and optimized data paths for high-throughput trade lifecycle management.
- Apply strong risk acumen to change management, capacity planning, and resilience testing (chaos engineering, failover, and BCP drills).
- Partner with engineers, domain experts, and global stakeholders to understand production processes, challenge entrenched assumptions in a cloud-centric, AI-driven world, and drive modernization.
- Multiply your impact with a modern, AI-centric toolchain, orchestrating AI agents across the SDLC and operations to rapidly comprehend large codebases, generate production-quality automation, and accelerate delivery.
Basic Qualifications
- 8+ years of professional software / reliability engineering experience, with stron
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s