Site Reliability Engineering Manager (Software Engineering - Mainframe)
FISAbout the role
Job Description
We are FIS. Our technology powers the world’s economy and our teams bring innovation to life. We champion diversity to deliver the best products and solutions for our colleagues, clients and communities. If you’re ready to start learning, growing and making an impact with a career in fintech, we’d like to know: Are you FIS?
NOTE: This position is hybrid (3 days onsite) at our FIS Office locations in Jacksonville (FL) & Milwaukee (WI).
About the Role and Team:
We are seeking a highly qualified Site Reliability Engineering Manager to lead operations within our fintech data center. The ideal candidate will bring deep expertise in SaaS platform reliability, server-side application management, Site Reliability Engineering principles, and demonstrate outstanding leadership abilities. Must have a proven track record in incident resolution, compliance-driven service delivery, and managing complex infrastructure and cross-functional teams in the fintech sector.
This individual will play a key role in driving the modernizing of critical applications with a focus on improving observability, automation, and resiliency. This is an opportunity to lead a team that will work across both mainframe technologies (COBOL, RPG) and modern server-based environments (Java, Angular, .NET), giving you a unique opportunity to operate at the intersection of legacy systems and contemporary microservices. This is a great opportunity to drive engineering improvements that directly enhance production support operations.
This individual will be responsible for ensuring environment reliability, toil automation, and resiliency improvements through effective oversight. Key responsibilities include strategic planning, team leadership, and fostering collaboration between technical and business units to ensure that operational initiatives align with organizational objectives. This role requires regular customer interactions regarding application performance, stability and reliability, including optimization and improved delivery of applications and services. This role is with our IBS Core Banking team.
What you will be doing:
Site Reliability Engineering Management: Oversee a team of Site Reliability engineers responsible for Identifying automation opportunities and implement tools and processes that streamline routine tasks, enable scalable infrastructure, and support seamless deployments.
Service Reliability and Availability Management: Lead improvement of the reliability and availability of critical applications, platforms, and server infrastructure through proactive monitoring, incident management, and resiliency improvements. Guide the team to develop and track new service level indicators to support SLO and SLA compliance.
Monitoring: Evaluate and interpret monitoring and alerting solutions that improve visibility into infrastructure, application performance, and user experience. Proactively identifying improvement opportunities and implementing effective corrective actions.
Strategic Planning: Formulate and execute strategic initiatives to enhance efficiency, including capacity planning, disaster recovery, and business continuity measures.
Disaster Recovery: Recommend and implement improvements to disaster recovery plans, backup strategies, and failover mechanisms.
Compliance: Ensure ongoing compliance with industry regulations, standards, and best practices, particularly in data security and privacy.
Innovation: Maintain up-to-date knowledge of emerging technologies and trends in Site Reliability Engineering, SaaS platform server management and fintech to drive continuous innovation within the team.
Infrastructure Management: Supervise maintenance, configuration, and reliability of all data center infrastructure, including servers, networks, and storage systems. Delivers a production server operations environment that meets all service level agreements, processing service level objectives, response time targets, and availability targets.
Security and Compliance: Oversee data security protocols and maintain adherence to regulatory and industry standards.
Incident Management:
Lead incident management processes, ensuring rapid resolution and clear communication with stakeholders.
Identify and drive improvements in reliability, performance, and efficiency through data and root cause analysis.
Participate in an on-call rotation to support critical production incidents. You’ll join a globally dis
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s