Lead Data Engineer
SupabaseAbout the role
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
We're looking for a seasoned Lead Data Engineer to join our growing data engineering team. You'll be responsible for leading the development and maintenance of our data infrastructure that powers analytics, machine learning, and operational AI workflows across the company. In this role, you will not only build critical systems but also guide the team's technical direction, mentor other engineers, and ensure the successful delivery of data projects. Your work will directly impact both our internal operations and our developer community through data-driven insights and products.
We believe in giving engineers the autonomy to work efficiently while setting high-performance standards. As we scale rapidly, we're seeking a leader who shares our commitment to open source, knows how to ship impactful features, and will contribute to and elevate our developer-focused culture.
The Stack
Data Warehouse: BigQuery (primary), PostgreSQL
Orchestration: Apache Airflow (Cloud Composer), Meltano
Data Modeling: dbt
Infrastructure: Google Cloud Platform, Pulumi
Analytics: Hex, PostHog
Reverse ETL: Hightouch
Languages: Python, SQL
What you'll own
Leadership & Project Management
Mentor and guide mid-level data engineers, fostering their technical and professional growth through code reviews and 1:1 guidance.
Lead the technical design and execution of complex data projects, ensuring alignment with architectural best practices and business objectives.
Partner with stakeholders across Analytics, Growth, and Engineering to define the team's roadmap, gather requirements, and prioritize data engineering initiatives.
Manage the team's backlog and workload, ensuring timely delivery of high-quality data pipelines and products.
Establish and evangelize best practices for data engineering, code quality, and documentation within the team.
Data Pipeline Development & Maintenance
Design, build, and maintain scalable ETL/ELT pipelines using Airflow and Meltano.
Develop and optimize dbt models following our established data warehouse architecture.
Implement data quality monitoring, testing, and alerting across all pipelines.
Manage data ingestion from 15+ sources including GitHub, HubSpot, Stripe, PostHog, Sentry, and internal PostgreSQL databases.
Infrastructure & Operations
Manage BigQuery datasets, reservations, and slot allocation across dev/staging/prod environments.
Deploy and maintain Airflow DAGs using Cloud Composer with custom Docker images.
Implement infrastructure as code using Pulumi for GCP resources.
Monitor pipeline performance and optimize for cost and efficiency.
Data Architecture & Modeling
Drive the evolution of our multi-layered data warehouse architecture with standardized naming conventions.
Design and implement data governance policies and documentation standards.
Optimize BigQuery performance through partitioning, clustering, and materialization strategies.
Reverse ETL & Data Activation
Develop and maintain reverse ETL pipelines to sync data to HubSpot, Customer.io, and other downstream systems.
Build attribution models and customer journey analytics.
Create automated triggers for sales and marketing outreach based on data milestones.
What you bring
Required Skills
5+ years of production experience with Python and SQL.
Proven experience mentoring other engineers and leading complex data projects from inception to completion.
Deep expertise in dbt for data modeling and transformation.
Extensive experience designing, deploying, and managing complex workflows in Apache Airflow.
Proficiency with cloud da
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s