Lead Observability Engineer, Global
Vantage Data CentersAbout the role
About Vantage Data Centers
Vantage Data Centers powers, cools, protects and connects the technology of the world’s well-known hyperscalers, cloud providers and large enterprises. Developing and operating across North America, EMEA and Asia Pacific, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
IT Standards Team
Our team is responsible for helping other technology teams on their automation journey, and for developing IT standards that support the IT organization. We embrace many approaches and technologies to speed up the delivery and operations of our Data Centers. From Zero-Touch provisioning of network equipment to the deployment of applications on containerization platforms, we apply our software and operation industry expertise everywhere we can. We question the status-quo and are not afraid to suggest new ways to do things. Individual contributors are encouraged to speak up, propose new insights and take an active role in the definition of our roadmap.
Position Overview
This role will be based remotely in the US.
Our team builds and operates the observability platform for Vantage Data Centers, enabling engineering and operations teams to understand system health, performance, and availability across data center and hybrid environments. To support our growth, we are looking for an experienced Observability Engineer with deep hands-on expertise in Elastic/Elasticsearch, Logstash, and Kibana, and a strong background creating and operationalizing metrics.
In this role, you will design, implement, and maintain end-to-end observability for logs and metrics: building resilient ingestion pipelines, defining schemas and parsing standards, creating Kibana dashboards and alerting, and partnering with platform, network, and application teams to set SLIs/SLOs and improve operational outcomes. You will continuously improve performance, reliability, retention, and cost of our telemetry pipelines while applying automation and infrastructure-as-code practices to keep the platform consistent and auditable.
Essential Job Functions
Design and operate a scalable observability platform with a primary focus on the Elastic Stack (Elasticsearch, Logstash, Kibana)
Build and maintain log ingestion and enrichment pipelines (Logstash) including parsing, normalization, and routing standards
Create, curate, and govern Kibana assets (dashboards, visualizations, Lens, Discover views) that support operations and engineering use cases
Define and implement metrics and alerting standards (SLIs/SLOs, thresholds, burn-rate alerts) to improve detection and reduce MTTR
Develop observability metrics for Operational Technology (OT) environments (e.g., BMS/EPMS/SCADA and other OT telemetry) to track availability, performance, alarms, and operational KPIs
Partner across teams to instrument services and infrastructure, troubleshoot incidents using telemetry, and continuously improve reliability, performance, and cost
Engineer and operate Elasticsearch clusters (or Elastic Cloud) including sizing, scaling, sharding/ILM, retention, backup/restore, and performance tuning
Develop and maintain Logstash pipelines (inputs/filters/outputs) to ingest logs/metrics from servers, network devices, virtualization platforms, containers, and cloud services
Create telemetry standards: field naming conventions, ECS alignment where appropriate, parsing/grok patterns, enrichment lookups, and data quality checks
Build and maintain Kibana dashboards, visualizations, and alerting rules; publish curated views for NOC/operations and engineering teams
Create and operationalize metrics: define SLIs/SLOs, implement metric collection/export, and ensure actionable alerting with runbooks and escalation paths
/Partner with facilities/critical infrastructure teams to define OT-focused SLIs/SLOs and metrics (e.g., alarm rates, sensor health, control loop status, device/point availability), normalize and tag OT telemetry, and build dashboards/alerts that support 24x7 operations
Automate configuration and deployment of observability components using infrastructure-as-code and configuration management (e.g., Terraform, Ansible) and CI/CD pipelines
Implement security best practices for telemetry platforms including role-based access control, data handling/PII controls, encryption, and auditability
Participate in incident response and post-incident reviews; use logs and metrics to identify root cause, document findings, and drive preventive improv
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s
Similar roles
Group Lead - Boxline
Oshkosh Corporation
Group Lead -Materials - 1st Shift
Oshkosh Corporation
Network Operations Team Lead
General Dynamics Information Technology
$138,000/yr