Jobs and Careers
BL
Senior Software Engineer / SRE - Trade Automation & Execution
BloombergNew York City, United Statesfull_timeVerifiedPosted 3 Feb 2026
💰 $240,000/yr($160,000/yr – $240,000/yr)
About the role
Senior Software Engineer / SRE - Trade Automation & Execution
Location
New York
Business Area
Engineering and CTO
Ref #
10049048
We are looking for an experienced engineer to help ensure that our real-time trading systems can scale safely, perform reliably under extreme market conditions, and recover gracefully from failures before issues impact clients.
Our Team:The TRAX Reliability team partners closely with application and infrastructure engineers to embed scalability, resilience, and technical risk management into trading systems from the ground up.
Rather than reacting to production incidents, we take a data-driven, proactive approach. We study real production workloads and controlled experiments to understand how systems behave under load, how failures propagate, and where bottlenecks emerge. By connecting performance, capacity, and risk, we help teams plan for growth, traffic spikes, and adverse scenarios with clear scaling strategies and recovery expectations.
We also design and build tooling that continuously evaluates system risk and performance. This includes running targeted stress tests, collecting detailed metrics, and surfacing insights through real-time dashboards. These tools enable teams to quickly identify bottlenecks across services, queues, and infrastructure, and to understand their impact on client experience.
What’s in it for you:
We’ll trust you to:
You’ll need to have:
Description & Requirements
The Trade Automation & Execution (TRAX) group builds the platforms and services that power modern electronic trading at Bloomberg. We design and operate high-performance, distributed, real-time systems used by financial institutions worldwide to execute trades, automate workflows, and make data-driven decisions. As markets evolve toward automation, scale, and intelligence, ensuring these platforms remain scalable, resilient, and predictable is critical. This is where TRAX Reliability plays a key role.We are looking for an experienced engineer to help ensure that our real-time trading systems can scale safely, perform reliably under extreme market conditions, and recover gracefully from failures before issues impact clients.
Our Team:The TRAX Reliability team partners closely with application and infrastructure engineers to embed scalability, resilience, and technical risk management into trading systems from the ground up.
Rather than reacting to production incidents, we take a data-driven, proactive approach. We study real production workloads and controlled experiments to understand how systems behave under load, how failures propagate, and where bottlenecks emerge. By connecting performance, capacity, and risk, we help teams plan for growth, traffic spikes, and adverse scenarios with clear scaling strategies and recovery expectations.
We also design and build tooling that continuously evaluates system risk and performance. This includes running targeted stress tests, collecting detailed metrics, and surfacing insights through real-time dashboards. These tools enable teams to quickly identify bottlenecks across services, queues, and infrastructure, and to understand their impact on client experience.
What’s in it for you:
- Have direct impact on the stability and resilience of execution platforms relied upon by the world’s leading buy-side firms
- Work on real-world, high-stakes distributed systems that need to operate under high performance and reliability requirements
- Develop deep expertise in scaling, failure modes, and technical risk management for real-time trading systems
- Collaborate with engineers across New York, London, and Frankfurt, significantly expanding your technical network
- Partner with application, observability, and infrastructure teams to influence system design across the organization
We’ll trust you to:
- Identify, prioritize, and track scalability and reliability risks across large-scale trading platforms
- Partner with application teams to diagnose and address performance and resilience challenges
- Analyze system behavior under real and simulated load, including latency, throughput, failover, and blast radius
- Design and run chaos engineering experiments and game-day exercises to validate system capacity and resilience
- Build and maintain automation and tooling for early detection and mitigation of production risks
- Communicate technical trade-offs, solutions, and roadmaps clearly to engineering stakeholders
- Plan for traffic growth and peak market events with clear scaling strategies and guardrails
You’ll need to have:
- 5+ years of professional experience with a high-level programming language such as Python, Java, or C++, preferably on Unix/Linux
- Solid understanding of Unix/Linux fundamentals
- Hands-on experience contributing to or triaging scaling and reliability issues in production distributed systems
- Experience working with metrics, monitoring, or observability platforms, such as Grafana, Prometheus, or log analytics tools
- Strong analytical skills and the ability to reason ab
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s