Jobs and Careers
AP

Incident Practice Manager

Apex Fintech Solutions
United Statesfull_timeVerifiedPosted 22 Aug 2025

About the role

WHO WE ARE

Apex Fintech Solutions (AFS) powers innovation and the future of digital wealth management by processing millions of transactions daily, to simplify, automate, and facilitate access to financial markets for all. Our robust suite of fintech solutions enables us to support clients such as Stash, Betterment, SoFi, and Webull, and more than 20 million of our clients' customers. 

Collectively, AFS creates an environment in which companies with the biggest ideas in fintech are empowered to change the world. As a global organization, we have offices in Austin, Dallas, Chicago, New York, Portland, Belfast, and Manila.

If you are seeking a fast-paced and entrepreneurial environment where you'll have the opportunity to make an immediate impact, and you have the guts to change everything, this is the place for you. 

AFS has received a number of prestigious industry awards, including:

  • 2021, 2020, 2019, and 2018 Best Wealth Management Company - presented by Fintech Breakthrough Awards

  • 2021 Most Innovative Companies - presented by Fast Company

  • 2021 Best API & Best Trading Technology - presented by Global Fintech Awards

ABOUT THIS ROLE

The Incident Practice Manager is a critical member of our Platform organization, specializing in both rapid incident response and proactive measures to ensure service reliability and minimize downtime. This role focuses on processes and practices of handling service disruptions, effective incident handling, and continuously improving incident processes to maintain a high standard of service availability and resilience. 

Duties/Responsibilities

  • Conduct ongoing training and education sessions; Ensure firm-wide understanding and adherence to the Incident Management processes and procedures 

  • Utilize and improve automation tools and scripts to streamline incident response and remediation, ensuring rapid and consistent handling of common problems. 

  • Lead practices for how to conduct effective RCA investigations; Guide teams on how to identify underlying or systemic causes, and document lessons learned to prevent recurrence. 

  • Implement incident response practices such as game-day exercise or chaos engineering to test and address our response to potential failures before they affect production 

  • Assess when and how to escalate incidents, ensuring the appropriate responders, teams, owners and stakeholders and are engaged for effective resolution. 

  • Facilitate comprehensive post-incident reviews, compiling insights and recommendations for process, tooling, or systems improvements. 

  • Develop, document, and regularly update incident management procedures and playbooks to support continuous improvement and onboarding of new staff. 

 

Education and/or Experience

  • Bachelor's degree in Computer Science, Information Systems, or a related field; advanced degree is a plus

  • 5+ years' experience leading incident investigations, root cause analysis and remediations 

  • 2+ years of experience as a people manager

  • Experience as an SRE or working extensively in an SRE environment preferred

  • Proven experience with relevant automation tools (e.g., scripting, runbooks, orchestration platforms) used in SRE and incident response. 

  • Extensive experience with ITSM, ITOM (e.g., ServiceNow, Jira Service Management) and Observability tools (PagerDuty, OpsGenie, Datadog, Prometheus, Grafana, etc.) preferred

  • Certification or training in incident management or IT service management (e.g., ITIL, ICS)  preferred

 

Required Skills/Abilities

  • Deep understanding of cloud systems, microservices, distributed systems, and the technical landscape of the organization. 

  • Exceptional analytical skills with the ability to quickly diagnose and resolve complex technical issues in high-pressure environments. 

  • Excellent verbal and written communication skills, enabling you to inform and coordinate across technical and non-technical stakeholders. 

  • Hands-on familiarity with incident management methodologies and frameworks, such as the Incident Command System (ICS). 

  • Strong knowledge of monitoring, logging, and tracing best practices to measure and interpret system behavior and performance.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Apex Fintech Solutions

View company profile →