Jobs and Careers
AS

Vice President, Head of Infrastructure Resiliency

AssetMark
United Statesfull_timeVerifiedPosted 29 Apr 2026
💰 $300,000/yr($250,000/yr – $300,000/yr)

About the role

Job Description:

As the Head of Platform Resiliency & Operations, you are accountable for operating and engineering the reliability, scalability, and resilience of AssetMark’s platform.

This role owns production operations today—including environments, batch processing, incident response, and day-to-day platform management—which are currently operationally intensive. Your mandate is to transform this reality by driving an engineering-first approach to production management and infrastructure.

You will lead a fundamental shift: from reactive, manual operations to proactive, automated, and engineered reliability—while continuing to deliver a high-quality, always-on platform for our clients.

This role has a twofold mandate:

Deliver on our client commitment by operating a high-availability, high-resiliency platform where reliability is a defining feature of the product Enable high-velocity product development by building systems, tooling, and practices that allow Product & Engineering to move fast without compromising stability

We can only consider candidates for this position who are able to accommodate a hybrid work schedule and are close to our Charlotte, NC office.

What You Will Own

1. Production Operations & Reliability Transformation

  • Own 24/7 production operations for mission-critical systems, including incident management, batch processing, and environment stability
  • Lead the transformation of production operations from manual, reactive processes to automated, engineering-driven systems
  • Establish an engineering-first mandate to eliminate manual toil and operational overhead
  • Drive systematic improvements in reliability, scalability, and operational efficiency

2. Reliability Engineering & Error Budget Management

  • Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across all critical systems
  • Establish and govern Error Budgets to balance product velocity with platform stability
  • Drive measurable reduction in operational toil through automation and engineering solutions
  • Embed reliability targets into planning and decision-making across teams
    • Apply Site Reliability Engineering (SRE) principles to quantify and manage reliability

3. Observability & Resilience Engineering

  • Build full-stack observability (metrics, logs, traces) to improve detection and diagnosis of issues
  • Evolve monitoring into deep observability with actionable alerting and reduced alert fatigue
  • Establish resilience testing practices (e.g., game days, fault injection)
  • Drive automated incident response and self-healing systems
  • Institutionalize blameless post-mortems focused on systemic improvement
  • Leverage SRE practices for incident learning and continuous improvement

4. Platform Engineering & Infrastructure

  • Ensure all infrastructure is managed via Infrastructure as Code (IaC) for consistency, scalability, and recovery
  • Own reliability and operational integrity of CI/CD pipelines, including automated release gating
  • Build self-service platforms and tooling that enable engineering teams to deploy and operate services safely
  • Modernize batch processing and environment management through automation and engineering rigor

5. Shared Reliability Ownership with Engineering

  • Establish shared accountability for reliability between Platform, SRE, and Software Engineering teams
  • Partner with Engineering to co-deliver reliability improvements and conduct joint post-incident reviews
  • Influence engineering practices including production readiness, safe deployments, and observability standards
  • Ensure reliability is embedded early in the software development lifecycle

6. Ecosystem & Vendor Reliability

  • Define and enforce reliability standards for third-party vendors and platform dependencies
  • Establish SLIs/SLOs for external services and manage vendor performance accordingly
  • Map and govern system dependencies to prevent cascading failures

How Success Is Measured

  • Sustained improvement in platform reliability as measured by SLO attainment
  • High availability and resiliency of client-facing systems
  • Reduction in operational toil and manual intervention across teams
  • Increased deployment velocity without degradation of reliability
  • Adoption of Infrastructure as Code and self-service platform capabilities
  • Reduction in incident frequency and improved detection (MTTD) and recovery (MTTR)
  • Demonstrated transformation from m

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

AssetMark

View company profile →