Jobs and Careers
SE

Principal Observability Architect

ServiceNow
United Statesfull_timeVerifiedPosted 2 Sept 2025

About the role

Company Description

It all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

Job Description

We are seeking a Principal Observability Architect to lead the strategic architecture, evolution, and operationalization of a modern, multi-tenant Observability Platform-as-a-Service (OPaaS) tailored for a hybrid on-prem and cloud-native SaaS product.

You will architect a cloud-agnostic, federated observability platform that supports real-time monitoring, advanced telemetry pipelines, and AI-powered insights to ensure platform reliability, developer productivity, and exceptional customer experiences. This role combines deep technical leadership with a strong focus on developer enablement, platform resiliency, and data governance.

What you get to do in this role:

Platform Architecture & Strategy

  • Lead architecture and roadmap for a multi-region, multi-cloud, multi-tenant observability platform scalable across diverse customer environments and service boundaries.
  • Architect near real-time telemetry ingestion pipelines with low-latency guarantees (seconds) using a mix of streaming and batch processing technologies.
  • Define observability blueprints including telemetry SLAs, data contracts, tenant data isolation, and cost-aware retention strategies for high-cardinality data.
  • Ensure observability systems are cloud-native and container-aware, supporting environments built on Kubernetes, service meshes, and serverless components.

Real-Time Monitoring & Detection

  • Design and implement real-time metrics, logs, traces, and event pipelines with technologies such as:
    • VictoriaMetrics, Prometheus, Grafana, Alertmanager
    • Cribl Stream and Edge for dynamic routing and filtering
    • VictoriaLogs for structured log analysis
  • Embed real-time anomaly detection and signal correlation, with context-aware alerting to reduce noise and MTTR.
  • Integrate with alerting and incident response tools (PagerDuty, Slack, ServiceNow) for automated incident routing and contextual enrichment.
  • Ensure observability of synthetic probes, end-user transactions, and critical SLOs with per-tenant granularity.

Instrumentation, Developer Enablement & CI/CD Integration

  • Standardize OpenTelemetry instrumentation across all services with prebuilt SDKs, language libraries, and semantic conventions.
  • Architect OpenTelemetry deployment patterns (agent-based, sidecar, collector pipelines) with support for Kubernetes, Lambda, and edge environments.
  • Embed observability validation gates into CI/CD workflows (e.g., GitHub Actions, GitLab CI) to enforce telemetry compliance before production rollout.
  • Provide self-service tools, templates, and training to enable developer teams to adopt observability by default.

AI for Observability & Productivity

  • Leverage AI/ML for:
    • Real-time anomaly detection and noise suppression
    • Predictive incident detection and impact forecasting
    • Auto-summarization of alert storms and telemetry bursts
    • Multi-tenant root cause and blast radius correlation
  • Build or integrate LLM-powered tools that support:
    • Natural language querying of live telemetry
    • AI-assisted debugging and dashboard generation
    • Generative runbooks and incident summaries

Data Platform Architecture

  • Architect hot and cold telemetry storage pipelines using:
    • VictoriaMetrics and Cribl for hot-path observability
    • Long-term retention in object storage (e.g., S3, GCS) using open formats (Parquet, JSON)
    • Federated querying engines like Trino for historical and cross-service analytics
  • Implement cost-aware ETL strategies, balancing real-time visibility with storage and ingestion optimization.
  • Incorporate data governance, PII handling, and regional data compliance (e.g., GDPR, SOC2) into telemetry architecture.

SaaS Operations & ITSM Integration

  • Integrate observability into ITSM and incident response systems (e.g., ServiceNow, Jira):
    • Auto-create incidents enriched with correlated traces, logs, and metrics
    • Provide real-time telemetry context

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

ServiceNow

View company profile →