Incident Practice Manager
Apex Fintech SolutionsAbout the role
WHO WE ARE
Apex Fintech Solutions (AFS) powers innovation and the future of digital wealth management by processing millions of transactions daily, to simplify, automate, and facilitate access to financial markets for all. Our robust suite of fintech solutions enables us to support clients such as Stash, Betterment, SoFi, and Webull, and more than 20 million of our clients' customers.
Collectively, AFS creates an environment in which companies with the biggest ideas in fintech are empowered to change the world. As a global organization, we have offices in Austin, Dallas, Chicago, New York, Portland, Belfast, and Manila.
If you are seeking a fast-paced and entrepreneurial environment where you'll have the opportunity to make an immediate impact, and you have the guts to change everything, this is the place for you.
AFS has received a number of prestigious industry awards, including:
2021, 2020, 2019, and 2018 Best Wealth Management Company - presented by Fintech Breakthrough Awards
2021 Most Innovative Companies - presented by Fast Company
2021 Best API & Best Trading Technology - presented by Global Fintech Awards
ABOUT THIS ROLE
The Incident Practice Manager is a critical member of our Platform organization, specializing in both rapid incident response and proactive measures to ensure service reliability and minimize downtime. This role focuses on processes and practices of handling service disruptions, effective incident handling, and continuously improving incident processes to maintain a high standard of service availability and resilience.
Duties/Responsibilities
Conduct ongoing training and education sessions; Ensure firm-wide understanding and adherence to the Incident Management processes and procedures
Utilize and improve automation tools and scripts to streamline incident response and remediation, ensuring rapid and consistent handling of common problems.
Lead practices for how to conduct effective RCA investigations; Guide teams on how to identify underlying or systemic causes, and document lessons learned to prevent recurrence.
Implement incident response practices such as game-day exercise or chaos engineering to test and address our response to potential failures before they affect production
Assess when and how to escalate incidents, ensuring the appropriate responders, teams, owners and stakeholders and are engaged for effective resolution.
Facilitate comprehensive post-incident reviews, compiling insights and recommendations for process, tooling, or systems improvements.
Develop, document, and regularly update incident management procedures and playbooks to support continuous improvement and onboarding of new staff.
Education and/or Experience
Bachelor's degree in Computer Science, Information Systems, or a related field; advanced degree is a plus
5+ years' experience leading incident investigations, root cause analysis and remediations
2+ years of experience as a people manager
Experience as an SRE or working extensively in an SRE environment preferred
Proven experience with relevant automation tools (e.g., scripting, runbooks, orchestration platforms) used in SRE and incident response.
Extensive experience with ITSM, ITOM (e.g., ServiceNow, Jira Service Management) and Observability tools (PagerDuty, OpsGenie, Datadog, Prometheus, Grafana, etc.) preferred
Certification or training in incident management or IT service management (e.g., ITIL, ICS) preferred
Required Skills/Abilities
Deep understanding of cloud systems, microservices, distributed systems, and the technical landscape of the organization.
Exceptional analytical skills with the ability to quickly diagnose and resolve complex technical issues in high-pressure environments.
Excellent verbal and written communication skills, enabling you to inform and coordinate across technical and non-technical stakeholders.
Hands-on familiarity with incident management methodologies and frameworks, such as the Incident Command System (ICS).
Strong knowledge of monitoring, logging, and tracing best practices to measure and interpret system behavior and performance.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s