Jobs and Careers
NC

Senior Site Reliability Engineer – Unified Observability

NCR Voyix
United Statesfull_timeVerifiedPosted 20 Jul 2026

About the role

About NCR VOYIX

NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.

Position Overview

We are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer Unified Observability initiative. This strategic role will be responsible for building and evolving a unified enterprise observability platform that delivers end-to-end visibility across NCR Voyix Restaurants, Retail, and Payments environments.

The ideal candidate will bring 10+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or related disciplines, with a proven track record of driving enterprise-scale observability, reliability, and operational excellence. This individual must be comfortable operating across organizational boundaries and partnering closely with Product Engineering, Infrastructure, Security, Operations, Architecture, and Executive Leadership teams to establish a comprehensive observability strategy and improve platform resilience.

This role will serve as a key technical leader responsible for defining standards, influencing architecture decisions, and enabling proactive operations through unified monitoring, telemetry, automation, and AI-driven insights.

Key Responsibilities

  • Lead the architecture, design, implementation, and continuous improvement of enterprise observability solutions across Azure, Google Cloud Platform (GCP), Kubernetes, and hybrid environments.
  • Establish and drive enterprise observability standards for monitoring, logging, distributed tracing, telemetry, and operational analytics.
  • Develop and maintain executive, operational, and engineering dashboards that provide real-time visibility into infrastructure, applications, platform health, customer experience, and business transactions.
  • Define, evangelize, and implement reliability frameworks including SLIs, SLOs, error budgets, operational KPIs, and service health metrics.
  • Partner cross-functionally with Engineering, Infrastructure, Security, Product, and Operations teams to identify reliability risks and drive operational excellence initiatives.
  • Lead efforts to improve incident prevention, detection, response, and recovery through intelligent alerting, automation, event correlation, and observability best practices.
  • Integrate observability capabilities with ServiceNow, CI/CD pipelines, automation frameworks, and enterprise operational workflows.
  • Influence technical strategy and roadmap decisions related to reliability engineering, platform observability, and operational readiness.
  • Support and drive enterprise initiatives involving AI-driven observability, predictive analytics, anomaly detection, and event intelligence.
  • Mentor engineers and serve as a subject matter expert for observability, reliability engineering, and cloud-native operations.
  • Establish governance, adoption, and best practices across multiple product and engineering teams to ensure consistent observability standards enterprise-wide.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 10+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related technical disciplines.
  • Demonstrated success designing and operating observability platforms in large-scale enterprise environments.
  • Deep expertise with Kubernetes platforms, including AKS and GKE.
  • Strong experience with Azure and Google Cloud Platform services and architectures.
  • Hands-on experience with enterprise observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, or similar platforms.
  • Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles.
  • Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks.
  • Strong automation and Infrastructure as Code expertise using Terraform and related tools.
  • Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages.
  • Experience integrating observability solutions into CI/CD pipelines and modern De

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

NCR Voyix

View company profile →