Vice President, Head of Infrastructure Resiliency
AssetMarkAbout the role
Job Description:
As the Head of Platform Resiliency & Operations, you are accountable for operating and engineering the reliability, scalability, and resilience of AssetMark’s platform.
This role owns production operations today—including environments, batch processing, incident response, and day-to-day platform management—which are currently operationally intensive. Your mandate is to transform this reality by driving an engineering-first approach to production management and infrastructure.
You will lead a fundamental shift: from reactive, manual operations to proactive, automated, and engineered reliability—while continuing to deliver a high-quality, always-on platform for our clients.
This role has a twofold mandate:
Deliver on our client commitment by operating a high-availability, high-resiliency platform where reliability is a defining feature of the product Enable high-velocity product development by building systems, tooling, and practices that allow Product & Engineering to move fast without compromising stability
We can only consider candidates for this position who are able to accommodate a hybrid work schedule and are close to our Charlotte, NC office.
What You Will Own
1. Production Operations & Reliability Transformation
- Own 24/7 production operations for mission-critical systems, including incident management, batch processing, and environment stability
- Lead the transformation of production operations from manual, reactive processes to automated, engineering-driven systems
- Establish an engineering-first mandate to eliminate manual toil and operational overhead
- Drive systematic improvements in reliability, scalability, and operational efficiency
2. Reliability Engineering & Error Budget Management
- Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across all critical systems
- Establish and govern Error Budgets to balance product velocity with platform stability
- Drive measurable reduction in operational toil through automation and engineering solutions
- Embed reliability targets into planning and decision-making across teams
- Apply Site Reliability Engineering (SRE) principles to quantify and manage reliability
3. Observability & Resilience Engineering
- Build full-stack observability (metrics, logs, traces) to improve detection and diagnosis of issues
- Evolve monitoring into deep observability with actionable alerting and reduced alert fatigue
- Establish resilience testing practices (e.g., game days, fault injection)
- Drive automated incident response and self-healing systems
- Institutionalize blameless post-mortems focused on systemic improvement
- Leverage SRE practices for incident learning and continuous improvement
4. Platform Engineering & Infrastructure
- Ensure all infrastructure is managed via Infrastructure as Code (IaC) for consistency, scalability, and recovery
- Own reliability and operational integrity of CI/CD pipelines, including automated release gating
- Build self-service platforms and tooling that enable engineering teams to deploy and operate services safely
- Modernize batch processing and environment management through automation and engineering rigor
5. Shared Reliability Ownership with Engineering
- Establish shared accountability for reliability between Platform, SRE, and Software Engineering teams
- Partner with Engineering to co-deliver reliability improvements and conduct joint post-incident reviews
- Influence engineering practices including production readiness, safe deployments, and observability standards
- Ensure reliability is embedded early in the software development lifecycle
6. Ecosystem & Vendor Reliability
- Define and enforce reliability standards for third-party vendors and platform dependencies
- Establish SLIs/SLOs for external services and manage vendor performance accordingly
- Map and govern system dependencies to prevent cascading failures
How Success Is Measured
- Sustained improvement in platform reliability as measured by SLO attainment
- High availability and resiliency of client-facing systems
- Reduction in operational toil and manual intervention across teams
- Increased deployment velocity without degradation of reliability
- Adoption of Infrastructure as Code and self-service platform capabilities
- Reduction in incident frequency and improved detection (MTTD) and recovery (MTTR)
- Demonstrated transformation from m
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s