Jobs and Careers
AL

Manager - Production Operations & Site Reliability Engineering

Alcon
Lake Forest, United Statesfull_timeVerifiedPosted 13 Aug 2026
💰 $181,500/yr($140,250/yr$181,500/yr)

About the role

Manager – Production Operations & Site Reliability Engineering

At Alcon, we are driven by the meaningful work we do to help people see brilliantly. We innovate boldly, champion progress, and act with speed as the global leader in eye care. Here, you’ll be recognized for your commitment and contributions and see your career like never before. Together, we go above and beyond to make an impact in the lives of our patients and customers.

We foster an inclusive culture and are looking for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability, availability, security, and continuous improvement of Alcon’s Digital Health Cloud platform supporting Software as a Medical Device (SaMD) and customer-facing digital health applications.

This role serves as the senior technical authority for production operations and Site Reliability Engineering (SRE), driving platform reliability, observability, automation, incident management, release governance, and cloud optimization. Partnering across Product Engineering, Architecture, Security, Infrastructure, and Operations, the Principal Engineer enables resilient, compliant, and scalable healthcare platforms with predictable, high-quality software delivery. In this role, a typical day will include:

Production Operations & Reliability

  • Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability, reliability, and performance.
  • Establish and continuously improve operational standards, SLIs/SLOs, readiness reviews, and service excellence practices.
  • Drive platform resilience, capacity planning, disaster recovery, and governance through Alcon’s Steady State Operations Framework (SSOF).

Release & Deployment Governance

  • Lead production readiness reviews, ensuring applications meet operational, security, monitoring, compliance, and supportability requirements.
  • Govern the Road to Production process across Validation, Staging, and Production environments, validating deployment readiness, infrastructure qualification, rollback plans, and acceptance criteria.
  • Partner with Engineering and DevOps teams to improve release quality, deployment reliability, and change success through standardized governance and automation.

Site Reliability Engineering (SRE)

  • Champion SRE best practices, leveraging automation and self-healing capabilities to reduce operational toil.
  • Improve service reliability and customer experience by reducing MTTD and MTTR and increasing deployment success rates.
  • Lead incident investigations, root cause analyses, and long-term corrective actions for critical production events.

Cloud Platform Operations

  • Provide technical leadership across AWS platforms including EKS, EC2, RDS, S3, ElastiCache, AWS MQ, Route53, and Kubernetes/Istio.
  • Optimize cloud infrastructure for scalability, resilience, security, performance, and cost efficiency.

Observability & Automation

  • Define enterprise observability strategies using Datadog, CloudWatch, distributed tracing, synthetic monitoring, centralized logging, and executive dashboards.
  • Lead automation initiatives across deployments, monitoring, health validation, incident response, and operational workflows.

Security & Compliance

  • Ensure compliance with HIPAA, GDPR, FDA, and enterprise cybersecurity standards.
  • Partner with Security teams to strengthen cloud security architecture, identity management, network segmentation, and operational controls.

Technical Leadership

  • Serve as the senior escalation point for major incidents, production events, and Hypercare operations.
  • Mentor engineering teams on operational excellence, production engineering, and SRE best practices.
  • Influence platform and architectural decisions that enhance operability, maintainability, resilience, and long-term service reliability.
  • Collaborate across Engineering, Architecture, Infrastructure, Security, and Global Operations to advance platform stability and operational maturity.

WHAT YOU’LL BRING TO ALCON:

Bachelor’s Degree or Equivalent years of directly related experience (or high school +13 yrs; Assoc.+9 yrs; M.S.+2 yrs; PhD+0 yrs)

The ability to fluently read, write, understand and communicate in English

5 Years of Relevant Experience

PREFERRED QUALIFICATIONS:

  • Bachelor’s Degree in degree in Computer Science, Engineering, or related field (Master’s preferred)
  • Experience in Production Operations, Site Reliability Engineering, Platform Engineering, or Cloud Operations
  • Experience supporting regulated healthcare or medical device platforms

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Alcon

View company profile →