Incident Manager / Vice President, Data Operations
OaktreeAbout the role
Oaktree is a leader among global investment managers specializing in alternative investments, with $202 billion in assets under management as of December 31, 2024. The firm emphasizes an opportunistic, value-oriented, and risk-controlled approach to investments in credit, private equity, real assets, and listed equities. The firm has over 1200 employees and offices in 23 cities worldwide.
Ā
We are committed to cultivating an environment that is collaborative, curious, inclusive and honors diversity of thought. Providing training and career development opportunities and emphasizing strong support for our local communities through philanthropic initiatives are essential to our culture.
Ā
For additional information please visit our website at www.oaktreecapital.com
Responsibilities
We are seeking a highly skilled and experienced Incident Manager to oversee the incident response lifecycle for a global asset management firm operating across multiple time zones. This role requires proactive incident handling, rapid problem resolution, and clear communication with technical and business stakeholders to minimize business disruptions. The Incident Manager will be responsible for enhancing and coordinating multi-system monitoring, escalation workflows, and continuous service improvement efforts. Based on lessons learned and industry best practices, the Incident Manager should drive initiatives to improve incident response processes, leverage tools and available documentation.
In addition, the role will involve working closely with various IT and business teams, ensuring cross-functional alignment, and driving automation and process optimization to enhance incident resolution efficiency. The individual in this role must be able to work under pressure, communicate extremely effectively, manage complex incidents, and facilitate decision-making at all levels of the organization. Of critical importance, the Incident Manager will lead post-incident reviews and contribute to strategic initiatives to strengthen operational resilience in the data operations space and mature incident response across the Oaktree organization.
Responsibilities include:
- Incident Lifecycle Management: Own and drive the end-to-end resolution of incidents affecting global data operations, while ensuring minimal downtime.
- Real-Time Monitoring & Detection: Utilize and leverage advanced monitoring tools (PagerDuty, ServiceNow, etc.) to detect system anomalies and trigger appropriate incident response measures.
- Incident Triage & Prioritization: Classify incidents based on severity (P1-P4), business impact, and urgency, ensuring a structured and efficient response
- War Room Activation & Coordination: Lead and facilitate war rooms for high-severity (P1/P2) incidents, engaging technical teams, business stakeholders, and executive leadership as needed.
- Communication & Stakeholder Management: Provide clear, timely updates to key stakeholders, including senior management, IT teams, and external partners, ensuring transparency and alignment. This is inclusive of initial incident communications outbound to impacted stakeholders, impact analysis and updates to all levels of the organization, and root cause analysis documentation outlining the path forward to a solution that will prevent a similar incident from taking place again.
- Escalation Management: Ensure proper escalation paths are followed, engaging engineering, infrastructure, cybersecurity, and third-party vendors when necessary.
- Root Cause Analysis (RCA) & Post-Incident Review: Lead post-mortem analysis for major incidents, document root causes, corrective actions, and make recommendations for long-term preventative measures.
- Process Optimization & Automation: Collaborate with all required stakeholders to implement automation, self-healing mechanisms, and runbook optimization.
- Regulatory & Compliance Reporting: Ensure incident response processes adhere to financial industry regulations (SEC, SOX, etc.) and collaborate with Compliance team members when required.
- Incident Response Metrics & Reporting: Track and report key performance indicators [Mean Time to Detect (MTTD), Mean Time to Acknowledge (MTTA), Mean Time to Resolve (MTTR), and incident recurrence rates to measure incident management effectiveness and drive continuous improvements.
- Training & Documentation: Develop and maintain incident management playbooks, escalation workflows, communication templates, and knowledge base articles to enhance response efficiency.
- Collaboration with Business & IT Teams: Work closely with business continuity, trading operations, and infrastructure teams to ensure cross-functional alignment and rapi
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights ā in under 60 seconds.
Apply Now āGenerate Application KitFree account required ā sign up in 30s