Jobs and Careers
OA

Infrastructure Engineering VR and DR Lead

Oaktree
Los Angeles, United Statesfull_timeVerifiedPosted 14 Sept 2025
💰 $200,000/yr($170,000/yr$200,000/yr)

About the role

Our Company

Oaktree is a leader among global investment managers specializing in alternative investments, with about $200 billion in assets under management. The firm emphasizes an opportunistic, value-oriented and risk-controlled approach to investments in credit, private equity, real assets and listed equities.  The firm has over 1,200 employees and offices in 24 cities worldwide.

We are committed to cultivating an environment that is collaborative, curious, inclusive and honors diversity of thought. Providing training and career development opportunities and emphasizing strong support for our local communities through philanthropic initiatives are essential to our culture.

For additional information please visit our website at www.oaktreecapital.com

Role Summary

The Infrastructure Engineering lead is a hands-on engineering leader that oversee the Virtualization and Data Center core infrastructure services with a focus on virtualization (VMware), data center operations, hardware lifecycle management, disaster recovery, and enterprise backup systems. Serving as a technical SME and escalation point for global operations, this role ensures stability, scalability, and security of infrastructure across Oaktree's on-premises and cloud environments.

This role requires expertise in managing hybrid computer environments, driving infrastructure standardization, improving operational maturity, and supporting critical business systems.

As a senior technical leader and escalation point, this position collaborates with teams across infrastructure, security, applications, and global site support to ensure consistent service delivery, audit readiness, and operational excellence.

Responsibilities

Virtualization and Compute Infrastructure

  • Design, implement, and maintain enterprise VMware environments (vCenter, ESXi, VM Tools).
  • Lead capacity planning, golden image management, disk provisioning, VM migrations, and host clustering.
  • Troubleshoot performance issues in virtualized environments; manage escalations related to server hangs and outages.
  • Support GPU-enabled virtualization environments for specialized workloads.

Data Center Operations

  • Lead operations for multiple data centers globally, including hardware deployments (Dell, Cisco UCS, VxRail), power/cooling management, and physical asset maintenance.
  • Oversee remote site infrastructure, coordinating with regional on-site support teams (e.g., in Stamford, New York, London, Singapore).
  • Provide on-the-ground support planning for AC replacements, hardware cabling, and equipment moves.

Asset and Configuration Management

  • Oversee ad-hoc asset tracking via spreadsheets and lead the migration toward ServiceNow CMDB and Lucidchart-based architecture documentation.
  • Drive standardization of asset documentation and support contract lifecycle visibility in ServiceNow ITAM (asset management)

Patch and Vulnerability Management

  • Coordinate quarterly patch (or as needed) cycles for physical and virtual infrastructure across global sites.
  • Implement vulnerability remediation in coordination with the cybersecurity team using SCCM, using Jira for tracking.
  • Ensure compatibility and test recovery post-patching to prevent disruptions.

Disaster Recovery (DR) and Business Continuity

  • Manage VMware Live Recovery Manager and Veeam-based DR platforms for failover and failback operations.
  • Execute DR tests across critical applications in coordination with the Business Continuity team.
  • Update and maintain DR runbooks and test documentation; support isolated bubble testing architecture.

Backup and Storage Management

  • Manage Dell IDPA/Data Domain backup appliances and policies.
  • Support NetApp and Dell PowerStore-based storage for SMB, NFS, and DFS shares; ensure proper failover to DR locations.
  • Oversee backup compliance for audit and legal hold (10-year archival); manage capacity thresholds and initiate storage expansion plans.

Site Support and Regional Oversight

  • Act as the escalation point and L3 support for major infrastructure incidents in Los Angeles and across remote offices.
  • Provide mentorship and architectural oversight to remote support leads (e.g., Phil in Stamford, Tony in Europe).
  • Support critical trading environments with dedicated infrastructure provisioning and incident response.

Documentation and Knowledge Management

  • Lead the effort

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Oaktree

View company profile →