Jobs and Careers
HG

Principal Data Engineer

HG Insights
Remotefull_timePosted 9 Sept 2026

About the role

<p><strong>About HG</strong><br>HG Insights is the pioneer of Revenue Growth Intelligence. For more than a decade, we have delivered comprehensive, AI- driven datasets on B2B buyers, technology adoption, IT spend, and buyer intent, sourced from billions of data points.<br>Today, we are a trusted partner to Fortune 500 technology companies, hyperscalers, and innovative B2B vendors seeking precise go-to-market analytics and decision-making. Through an evolving suite of AI agents that incorporate first-party data and buyer signals, HG Insights enables AI-powered GTM automation across sales, marketing, RevOps, and data analytics teams, modernizing GTM execution from strategy through activation.</p> <p><strong>About the Role</strong><br>This is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands. It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result.</p> <p>You will be based in Pune, working with engineers at our Pune, US &amp; Brazil locations.</p> <p><strong>What you’ll do</strong></p> <ul> <li>Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.</li> <li>Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring, and pipelines that quarantine bad data rather than publish it.</li> <li>Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.</li> <li>Make our release cycle boring. Recurring deliveries should not depend on people watching them.</li> <li>Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.</li> </ul> <p>&nbsp;</p> <p><strong>What we are looking for</strong></p> <ul> <li>15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.</li> <li>Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.</li> <li>Strong SQL and dimensional modelling, and the judgment to know when to break the rules.&nbsp;</li> <li>Experience with performant and scalable OLTP setups (MySQL/Postgres).</li> <li>Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databric

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

HG Insights

View company profile →