VP, Reliability Engineer - OnePay
SynchronyAbout the role
Job Description:
Role Summary / Purpose:
As the pace and complexity of software development accelerates, maintaining the reliability, performance, and scalability of critical financial systems is more vital than ever. OnePay, as a strategic partner, expects seamless integration across APIs, batch processes, file transmissions, and operational workflows with rigorous
SLAs, blazing speed, and superior customer experience.
The VP, Reliability Engineer for OnePay will lead efforts to embed reliability engineering principles across the full Synchrony technology stack that enables OnePay, ensuring high availability, fault tolerance, and rapid incident response. This leadership role bridges software engineering, system administration, and DevOps to build resilient, self-healing processes and platforms that meet and push boundaries beyond Synchrony & OnePay standards.
Your mission is to drive innovation and continuous improvement in monitoring, alerting, automation, and root cause analysis to reduce downtime, improve system resiliency, and enhance end-user satisfaction through flawless operational performance.
Our Way of Working
We’re proud to offer you choice and flexibility. At Synchrony, our way of working allows you to have the option to work from home, near one of our Hubs or come into one of our offices. Occasionally you may be required to commute to our nearest office for in person engagement activities such as business or team meetings, training and culture events.
Essential Responsibilities:
Lead reliability engineering initiatives for OnePay’s full integration landscape, including APIs, batch processes, file transmissions, and operational workflows, ensuring they meet or exceed SLAs and service reliability objectives (SLOs).
Design and implement scalable, self-healing systems that promote fault tolerance and rapid recovery, minimizing customer impact while supporting OnePay’s appetite for speed and responsiveness.
Collaborate closely with OnePay & Synchrony development teams, platform engineers, and product stakeholders to embed reliability best practices throughout the software lifecycle.
Establish rigorous monitoring, alerting, and real-time telemetry focusing on key KPIs that measure availability, latency, throughput, and error rates aligned with OnePay SLAs.
Own root cause analysis (RCA) and continuous improvement processes, systematically eliminating recurring issues and driving a culture of operational excellence and resilience.
Drive automation efforts to simplify operational complexity, including CI/CD pipeline integration, automated rollbacks, and canary deployments.
Coordinate cross-functional engagements between infrastructure, security, application teams, and Synchrony + OnePay business leaders to ensure unified reliability strategies and transparent communication.
Lead incident response efforts, including on-call leadership, rapid diagnostics, remediation, and stakeholder reporting to meet client expectations for responsiveness.
Influence and coach development teams on reliability design patterns such as graceful degradation, rate limiting, retry mechanisms, and service mesh integrations.
Stay current with fintech industry trends, OnePay-specific regulatory requirements, and emerging technology to anticipate and mitigate risk.
Facilitate & influence best practice sharing across our Reliability Engineering Community of Practice aimed to enhance enterprise & OnePay resiliency
Manage special projects and other duties as assigned.
Qualifications / Requirements:
Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience (minimum 5 years combined experience in software development and systems reliability/support) OR in Lieu of degree, 10+ years of experience within Software development and Systems reliability.
Minimum 5 years’ experience in full-stack development with technologies including Spring Framework, Java, REST APIs, and front-end UI frameworks.
Extensive experience designing, deploying, and troubleshooting large-scale distributed systems and service-oriented architectures, ideally within payment or fintech ecosystems.
Proven track record in systems engineering, reliability engineering, and infrastructure operations.
Experience influencing development teams in planning product implementations to tackle and resolve technical debt
Deep understanding of API gateways, batch processing workflows, file transmission protocols, and operational monitori
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s