Jobs and Careers
IN

Senior Director, Customer Reliability & Technical Operations (Pharmacy)

Inovalon
Remote- United States, United StatesRemotefull_timeVerifiedPosted 24 Mar 2026
💰 $260,000/yr($196,600/yr$260,000/yr)

About the role

Inovalon was founded in 1998 on the belief that technology, and data specifically, would empower the transformation of the entire healthcare ecosystem for the better, improving both outcomes and economics. At Inovalon, we believe that when our customers are successful in their missions, healthcare improves. Therefore, we focus on empowering them with data-driven solutions. And the momentum is building.

Together, as ONE Inovalon, we are a united force delivering solutions that address healthcare’s greatest needs. Through our mission-based culture of inclusion and innovation, our organization brings value not just to our customers, but to the millions of patients and members they serve.

Overview

The Senior Director, Customer Reliability & Technical Operations is accountable for the reliability, availability, performance, and operational health of Inovalon customer pharmacies using the ScriptMed SaaS platform. This leader directs a global operations organization (United States and India) responsible for ITIL-aligned Incident, Problem, and Change Management, as well as the technical functions that keep the platform stable and scalable, including Cloud Infrastructure Engineering, Database Administration, DevOps, and Site Reliability Engineering (SRE). The role partners closely with Product, Engineering, Security, and Customer Success to proactively detect and remediate issues using DataDog observability and ServiceNow ITSM workflows, ensuring customers experience dependable service and predictable outcomes.

Scope and Impact

Owns day-to-day and sustained operational performance for ScriptMed, including uptime, performance, incident response, and service restoration across customer pharmacies and tenant environments. Leads a blended onshore/offshore operating model, ensuring 24x7 coverage, clear escalation paths, and consistent execution of operational processes. Establishes and matures a Network Operations Center (NOC) and evolves it into an AI-enabled Intelligent Operations Management Center, improving detection, triage, and automation. Drives operational discipline across Incident, Problem, and Change Management, reducing customer-impacting events, shortening MTTR, and preventing recurrence. Provides executive-level visibility into platform health and risk, enabling informed decisions on investment, capacity, and reliability improvements.

Key Responsibilities

Leadership & Operating Model

Define and execute the customer reliability and technical operations strategy aligned to ScriptMed business objectives, SLAs, and customer expectations. Build and lead high-performing teams across the U.S. and India, including staffing, performance management, coaching, and career development. Establish clear on-call and escalation models, operational playbooks, and governance routines (daily ops review, incident review, weekly change review, reliability council). Partner with Engineering, Product, Security, and Customer Success to align priorities, manage operational risk, and drive continuous improvement.

Incident Management (Service Restoration)

Own the incident management lifecycle, including detection, triage, escalation, customer-impact assessment, communications, and service restoration. Ensure strong runbooks, incident roles, and standards for severity classification, timelines, and stakeholder updates. Use DataDog monitoring and alerting to proactively identify issues and reduce customer impact through early detection and fast response. Lead post-incident reviews, ensuring corrective actions are assigned, tracked, and validated.

Problem Management (Prevention and Root Cause)

Establish and mature a problem management program that drives root cause analysis, corrective and preventive actions, and measurable reduction in repeat incidents. Create a consistent approach for trend analysis, known error management, and prevention backlog creation. Partner with Engineering and Architecture to prioritize reliability improvements and reduce technical debt that drives operational instability.

Change Management (Risk Reduction)

Own change management processes to ensure reliable deployments, infrastructure changes, and operational updates with appropriate approvals and controls. Define change classification standards, risk scoring, blackout windows, validation steps, and rollback plans. Partner with DevOps and Engineering to implement change automation and quality gates that reduce change-related incidents.

NOC / Intelligent Operations Management Center

Stand up and operationalize a Network Operations Center responsible for real-time monitoring, initial triage, and coordinated response. Mature NOC capabilities into an AI-enabled Intelligent Operations Management Center that uses

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Inovalon

View company profile →