Enterprise Observability Architect
MSDAbout the role
Job Description
We are seeking a visionary and influential technical leader to join as an Enterprise Observability Architect. In this role, you will be the chief architect and evangelist for our company-wide observability strategy. Your mission is to define and drive the adoption of a unified, full-stack observability framework that empowers our product teams to build and operate highly reliable, performant, and resilient services.
This is not just a technical role; it is a strategic one. You will work across organizational boundaries, from engineering teams to senior leadership, to champion a culture of operational excellence. You will be responsible for rationalizing our tooling landscape, eliminating redundancy, and designing a cohesive, self-service observability platform that is easy to consume. If you are passionate about turning complex technical concepts into clear business value and leading large-scale change, this is your opportunity to make a lasting impact.
What You'll Do
Create Strategy: Create, document, and champion a unified, long-term observability strategy for the entire company, covering all telemetry types like metrics, traces, logs, profiles and more
Architect the Future: Design a cohesive, full-stack observability solution using best-in-class tools and practices. Ensure our architecture promotes a product-oriented approach with a strong focus on self-service capabilities
Drive Standardization and Consolidation: Actively identify and help decommission redundant or overlapping tools, driving the organization towards a standardized, cost-effective, and easy-to-manage observability service
Join forces with SRE team: Partner with Site Reliability Engineering (SRE) team, promote their principles, such as defining, measuring, and managing Service Level Objectives (SLOs) and error budgets and how critical Observability for them
Evangelize and Educate: Clearly and persuasively communicate complex observability topics, benefits, and strategic decisions to senior leadership and diverse stakeholders across IT and business domains
Lead by Influence: Act as the senior technical expert and thought leader in the observability space. Mentor engineers and guide teams in implementing best practices for instrumentation (OpenTelemetry), monitoring, and incident analysis using AIOps
Stay on the Cutting Edge: Continuously research industry trends, emerging technologies (like eBPF, OpenTelemetry), and best practices to ensure our observability strategy remains modern and effective
Must-Have Skills & Experience
Extensive Technical Experience: 10+ years in Observability, software engineering, SRE, or platform engineering, with a proven track record of architecting and implementing large-scale technical solutions
Deep Observability Expertise: Subject matter expert on Observability practices and collection of all types of telemetry. SME for the tools that support them (e.g., Grafana, Prometheus, Dynatrace, BigPanda, xMatters or others)
Strategic Leadership: Demonstrable experience creating and executing a technical strategy across multiple teams or an entire organization
SRE and SLO Mastery: Deep, practical knowledge of Site Reliability Engineering principles, with hands-on experience defining, implementing, and managing SLOs and error budgets
OpenTelemetry Proficiency: Strong understanding and practical experience with the OpenTelemetry standard for instrumentation and telemetry collection
Multi-Cloud Experience: Experience designing solutions that operate across multiple cloud providers (AWS, Azure, GCP)
Exceptional Communication & Influence: World-class ability to explain highly complex technical concepts to senior executives and non-technical audiences. Proven ability to lead by influence and drive consensus in a large organization
Nice-to-Have Skills & Experience
Tool Rationalization: Experience leading projects to consolidate tooling and migrate teams to a centralized platform
FinOps Knowledge: Understanding of cloud cost management principles and how they apply to observability data and tooling
eBPF Knowledge: Familiarity with eBPF for advanced, low-level system observability and performance analysis
Product Mindset: Experience thinking about internal platforms as products, with a focus on user experience, self-service, and clear documentation
Why Join Us?
Unmatched Impact: Define the technical direction for a critical capability ac
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s