Jobs and Careers
AL

Director, Infrastructure & Site Reliability Engineering

Alteryx
United StatesRemotefull_timeVerifiedPosted 13 Jul 2026
💰 $239,610/yr($181,900/yr$239,610/yr)

About the role

Meet the Moment with Alteryx

 

We're living through a once-in-a-generation shift in how work gets done. Data, automation, and AI are quickly becoming the center of every business decision - and Alteryx is leading the transformation.

 

You'll be working on the challenges that sit at the heart of modern business. No matter your role, the work you do will help organizations move faster, see more clearly, and tackle questions that used to feel impossible.

 

If you're ready to meet the moment with innovation, curiosity, and excellence, there's a place for you here.

Alteryx is searching for a Director, Infrastructure & Site Reliability Engineering. This position is remote-friendly.

Position Overview:

We are seeking an experienced engineering leader to lead Alteryx’s Infrastructure, Site Reliability Engineering (SRE), Observability, and Performance Engineering organizations. In this role, you will define the technical vision and execution strategy for the platforms and operational capabilities that power our cloud services, enabling engineering teams to build, deploy, and operate reliable, secure, and highly scalable products.

You will lead multiple engineering teams responsible for cloud infrastructure, reliability engineering, observability, performance optimization, and operational excellence. This leader will partner closely with Product Engineering, Security, Compliance, and Customer Operations to ensure our platform meets the highest standards for availability, scalability, security, and customer experience.

Primary Responsibilities:

  • Define and execute the strategy for Alteryx’s centralized Infrastructure, Site Reliability Engineering (SRE), Observability, and Performance Engineering organizations.

  • Lead the design, operation, and continuous evolution of cloud infrastructure across AWS and GCP, ensuring scalability, reliability, security, and cost efficiency.

  • Drive Infrastructure-as-Code adoption and governance through Terraform, establishing consistent platform standards, automation, and operational best practices.

  • Own the company’s observability strategy by building and operating enterprise-grade telemetry platforms using Datadog and related technologies, enabling actionable insights into system health, performance, and customer experience.

  • Partner with Security, Compliance, and Engineering teams to meet regulatory and customer requirements, including HIPAA, FedRAMP, SOC 2, and other compliance frameworks.

  • Establish and continuously improve incident management practices, including operational readiness, on-call excellence, postmortem culture, root cause analysis, and measurable reliability improvements.

  • Develop proactive reliability programs including capacity planning, resiliency testing, disaster recovery, performance benchmarking, and operational risk management.

  • Define reliability engineering frameworks that enable product teams to own service health through Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, performance objectives, and operational accountability.

  • Lead the evolution of centralized platform capabilities that simplify how engineering teams build, deploy, monitor, and operate services at scale.

  • Partner with engineering leadership to improve developer productivity through platform automation, self-service infrastructure, deployment tooling, and operational best practices.

  • Build, mentor, and develop high-performing engineering managers and technical leaders while fostering a culture of operational excellence, customer focus, accountability, continuous learning, and innovation.

Qualifications:

  • 10+ years of software engineering, infrastructure, or platform engineering experience, with 5+ years leading multiple engineering teams or managers.

  • Proven experience leading Infrastructure, SRE, Platform Engineering, or Cloud Operations organizations supporting large-scale SaaS products.

  • Deep expertise operating production environments on AWS and/or GCP.

  • Strong experience with Infrastructure-as-Code technologies such as Terraform.

  • Experience building and operating modern observability platforms using Datadog, OpenTelemetry, Prometheus, Grafana, or similar technologies.

  • Demonstrated success implementing SRE practices including SLOs, SLIs, error budgets, incident management, operational reviews, and reliability engineering programs.

  • Experience supporting regulated environments and working with compliance frameworks such as HIPAA, FedRAMP, SOC 2, ISO 27001, or similar.

  • Strong understanding of distributed syste

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Alteryx

View company profile →