Jobs and Careers
OM

Sr. Manager, Site Reliability

Omnicell
United Statesfull_timeVerifiedPosted 21 Jul 2026

About the role

 

About this opportunity

Omnicell is building a Global Cloud Operations organization from the ground up as our business shifts from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability engine of that organization, and this role is the first senior SRE hire — the person who will design the practice, set the standards, and then run the plays themselves until the team is large enough to delegate.

This is not a role where reliability practices already exist and you tune them. It is a role where you define what good looks like for Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud Operations and an Engineer III SRE you will coach and grow.

The environment is hybrid. Some of our products are still hardware in hospitals communicating with cloud services; others are fully SaaS. Some customers access us over private circuits, others over the public internet. We operate in a regulated environment — HIPAA, SOC 2, and in some engagements FedRAMP — which means reliability, security, and auditability are not separable concerns. The person we hire will be comfortable with that complexity and will help the organization design for it rather than around it.

This role also anchors Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner of how that gets introduced, prioritized against foundational reliability work, and validated in a regulated environment.

What you will own

Reliability practice (the coach half)

  • Define and publish SLOs and SLIs for the top 5–10 Tier-1 customer-facing services, in partnership with Product and Engineering. Establish error budget policy and the enforcement mechanism when budgets burn.

  • Design the incident command structure: severity rubric, declaration criteria, war-room protocol, stakeholder communication cadence, and the postmortem template. Train the first cohort of incident commanders across Engineering and Support.

  • Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet.

  • Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and Enterprise Security — into a durable SRE-owned model.

  • Establish the on-call rotation model, including fair distribution, compensation approach, paging discipline, and the handoff protocol with our existing managed services partners (IBM, HCL) who provide L1/L2 coverage.

  • Develop and track operational KPIs — MTTR, SLO attainment, change-failure rate, recurrence, cost per workload — and present reliability metrics and improvement roadmaps to senior leadership in the monthly Cloud Ops executive review.

Hands-on engineering (the player half)

  • Instrument Tier-1 services yourself. Write the dashboards. Write the alerts. Write the runbooks. Do not wait for the team to grow before the work starts.

  • Take the pager. Commander Sev-1 and Sev-2 incidents until a broader on-call rotation is staffed. Lead blameless postmortems and drive follow-up work to resolution.

  • Contribute code and infrastructure-as-code (Terraform preferred; Chef/Puppet acceptable) to the platform. Oversee the design and evolution of CI/CD pipelines — our current stack includes CodeFresh, TeamCity, GitHub Actions, and Octopus Deploy, and we are consolidating over time.

  • Administer and scale our Kubernetes platform, including secure and compliant cluster configurations. Working knowledge of Docker, Helm, and Service Mesh (Istio or Linkerd) expected.

  • Run chaos and failover exercises (Chaos Monkey, LitmusChaos, or equivalent). Validate that what we think is resilient actually is.

AI-driven operations

  • Architect Omnicell's AIOps direction: evaluate and introduce ML-based anomaly detection, predictive alerting, automated root cause analysis, and LLM-assisted runbook or triage pipelines.

  • Make informed build-versus-buy calls across the AIOps landscape. Integrate AI-assisted tool

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Omnicell

View company profile →