Jobs and Careers
DA

Staff Site Reliability Engineer (Databricks)

Datavant
Remote - United States, United StatesRemotefull_timeVerifiedPosted 25 Jun 2025
💰 $250,000/yr($200,000/yr$250,000/yr)

About the role

Datavant is a data platform company and the world’s leader in health data exchange. Our vision is that every healthcare decision is powered by the right data, at the right time, in the right format.

Our platform is powered by the largest, most diverse health data network in the U.S., enabling data to be secure, accessible and usable to inform better health decisions. Datavant is trusted by the world’s leading life sciences companies, government agencies, and those who deliver and pay for care. 

By joining Datavant today, you’re stepping onto a high-performing, values-driven team. Together, we’re rising to the challenge of tackling some of healthcare’s most complex problems with technology-forward solutions. Datavanters bring a diversity of professional, educational and life experiences to realize our bold vision for healthcare.

What We’re Looking For

We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission-critical data and ML workloads across our organization.

This role is ideal for someone who combines a strong SRE mindset with deep cloud infrastructure and data platform experience. You're comfortable operating at scale in a complex, hybrid cloud environment and can architect systems that balance velocity, safety, and cost. You’ll work closely with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to build a modern data platform that is secure, self-service, and production-grade.

What You Will Do

  • Operate and Improve Databricks: Own Databricks platform lifecycle—including automation, workspace governance, job orchestration, and cost optimization.

  • Design for Reliability: Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.

  • Advance Observability: Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.

  • Drive CI/CD for Data & ML: Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.

  • Enable Data Flow Across Platforms: Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.

  • Champion Event-Driven Architectures: Leverage cloud-native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.

  • Collaborate Across Teams: Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.

  • Contribute to Strategy: Influence engineering-wide decisions on data platform architecture, ML enablement, and data product strategy.

 

What You Need to Succeed

    • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.

    • Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools.

    • Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.

    • Proven expertise with observability tools (especially Datadog) and architecting platform-wide logging and monitoring solutions.

    • Strong command of CI/CD tooling, especially GitHub Actions, infrastructure-as-code (Terraform), and deployment automation for data systems.

    • Strong programming/scripting skills in Python, Bash, or Go.

    • Experience building and supporting highly available, fault-tolerant systems.

    • Excellent communication and collaboration skills; able to work effectively across teams.

 

Nice to Haves

  • DevSecOps mindset: Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.

  • Experience with ML infrastructure tooling<

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Datavant

View company profile →