Director, Site Reliability
Early WarningAbout the role
At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Engineering at Early Warning (EWS) is a blend of teams organized around many different platforms, capabilities and products that are brought together to power core capabilities at the biggest banks in America – this includes ubiquitous products like Zelle® and Paze.These capabilities are typically provided behind a customer-facing API or integration point which enables the EWS teams to innovate aggressively where big wins can be found. The teams aligned behind these efforts drive their own innovation in partnership with stakeholders.
If you are hungry for large scale challenges and crave opportunities to learn and contribute in a big way – we’d love to talk to you!
Overall Purpose
The Director, Site Reliability is a highly impactful role responsible for ensuring the reliability of EWS applications by managing the EWS Platform Site Reliability team. The SRE team supports our first responders in the tools and information they use, while driving mitigations for all high priority incidents and owning the RCA program to prevent incidents from recurring.
Essential Functions
Leads a high performing team of Site Reliability Engineers that provides 24/7 response to critical incidents.
Develops, manages, and owns tools that provide observability, monitoring and alerting for all EWS product applications.
Works closely with product and infrastructure development teams to ensure our applications are instrumented and measured.
Responsible for implementing all changes through pipelines and code.
Identify, evangelize and implement reliability patterns for EWS product applications.
Owns and improves the Incident Management Program, Policy, and Procedures for EWS, including both handling active incidents and the Root Cause Analyses process.
Owns the Change Management function at EWS ensuring changes to critical environments have risk appropriately identified and mitigated.
Supports the company’s commitment to risk management and protecting the integrity and confidentiality of systems and data.
The above job description is not intended to be an all-inclusive list of duties and standards of the position. Incumbents will follow instructions and perform other related duties as assigned by their supervisor.
Minimum Qualifications
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s