Jobs and Careers
KE

Director of Observability

Kestra Holdings
United Statesfull_timeVerifiedPosted 15 Jul 2026

About the role

Kestra Holdings offers industry-leading wealth management platforms for independent wealth management professionals nationwide. Kestra is dedicated to empowering independent financial professionals—including traditional and hybrid RIAs—to grow their businesses and deliver exceptional client service. We combine advanced business management technology with personalized consulting to provide unmatched scale, efficiency, and support. Our advisor-focused culture is built on innovation and advocacy, enabling advisors to offer comprehensive securities and investment advisory solutions to their clients.


Lead with Purpose. Partner with Impact.
We are seeking a Director of Observability to stand up a brand-new observability and reliability practice from the ground up at Kestra Holdings. This is a working leadership role — the Director will be expected to be hands-on in the early phase: selecting and configuring tooling, writing instrumentation standards, building the first dashboards and alerting pipelines, and personally running incident command for major events while the team and platform mature.
This is a newly created leadership role reporting directly to the Head of IT Infrastructure & Cybersecurity. The Director will start with two direct reports — a to be hired Senior Observability Architect (India-based) and a future US-based Observability/Reliability Engineer — and will be expected to scale the team over time as the practice and service catalog grow.

What you’ll Do:
• Observability Strategy & Platform Hands-On Build-Out.
• Define and execute the observability strategy for Kestra Holdings, aligned with business objectives, regulatory requirements, and the enterprise technology roadmap.
• Personally lead the initial build-out of the observability platform across metrics, logs, traces, profiles, and alerting — including tool evaluation, POCs, architecture, deployment, and configuration (e.g., Azure Monitor/Log Analytics, Datadog, Grafana, OpenTelemetry, Elastic/Splunk).
• Work with and enforce existing instrumentation standards (OpenTelemetry, structured logging, distributed tracing) across infrastructure and application teams.
• Build the first generation of dashboards, SLO scorecards, and a single pane of glass for Tier-1 service health — rolling up sleeves alongside the Sr. Architect and engineer.
• Operate the firm's end-to-end incident management lifecycle — detection, response, escalation, communication, and blameless post-incident review.
• Stand up on-call schedules, escalation policies, and runbook-driven triage for Sev1–Sev4 incidents via PagerDuty/xMatters or equivalent.
• Serve as primary incident commander for major incidents during the initial build phase, transitioning command responsibilities to senior team members as the practice matures.
• Integrate the incident lifecycle with Jira / Jira Service Management (JSM) for ticketing, change correlation, and remediation tracking; partner with Cybersecurity so incidents run once, not separately by Infra and Cyber
• Facilitate post-incident reviews (PIRs), track remediation items in Jira, and report trends to leadership.
• Drive adoption of SRE principles across the firm: SLI/SLO definition, error budget policy and enforcement, toil identification and automation, and operational readiness reviews.
• Establish release of reliability gates and embed reliability into the service lifecycle from design through production.
• Partner with Cloud & Platform Engineering, Cybersecurity, and application teams to ensure all services are fully instrumented, measurable, and integrated into the firm's SLO and incident frameworks.
• Ensure observability and incident management practices align with NIST CSF 2.0 maturity targets and support the firm's cybersecurity roadmap.
• Partner with the Cybersecurity team to integrate observability data with Jira/JSM, CMDB, and SIEM for enriched context during incidents.
• Support regulatory and audit requirements appropriate for a SEC-regulated financial services firm (e.g., logging retention, evidentiary integrity, access controls on telemetry data).
• Directly lead and mentor a small, high-leverage team of two to start: a Senior Observability Architect (India-based) and a US-based Observability/Reliability Engineer.
• Operate as a player-coach — splitting time between strategic leadership, hands-on engineering, and direct mentorship of the senior architect.
• Build a multi-year workforce plan and talent pipeline to scale the team as the platform, service catalog, and 24/7 coverage needs grow.
• Foster a culture of blameless learning, operational excellence, and engineering-led reliability across US and India hours of coverage.
• Serve as the primary technical liaison for observability and incident management vendors (e.g., Datadog, PagerDuty/

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Kestra Holdings

View company profile →