About the role
<h2><strong>About the Role </strong></h2> <p>We are looking for a Senior Backend Engineer to lead the unification of <strong>large, highly rich, and heterogeneous datasets</strong> sourced from a wide range of external providers. These datasets are used to power our generative audio models.&nbsp;</p> <p>Your work will create the foundational dataset that powers our research by building robust, scalable systems for <strong>linking, deduplicating, reconciling, and enriching </strong>data at massive scale. This role centers on&nbsp;<strong>high-impact bulk ingestion and advanced data linkage</strong>. You will design the logic, algorithms, and strategies that transform many independent datasets into a unified, high-quality canonical asset used throughout the company.</p> <p>You will collaborate closely with ML researchers and product teams, working with tools such as <strong>BigQuery, Dataflow/Beam, TFRecords</strong>, and—where beneficial—distributed systems frameworks like&nbsp;<strong>Ray</strong>. Familiarity with ML workflows using&nbsp;<strong>JAX</strong> or <strong>multihost training</strong> is a plus, as the datasets you produce will directly support that ecosystem.</p> <h2>What You'll Do</h2> <ul> <li>Build high-throughput&nbsp;<strong>bulk ingestion workflows</strong>&nbsp;to integrate datasets from multiple external providers.&nbsp;</li> <li>Design and implement scalable&nbsp;<strong>entity-resolution</strong>&nbsp;solutions, including record linking, deduplication, clustering, and conflict arbitration.&nbsp;</li> <li>Create and refine&nbsp;<strong>matching logic, decision rules, and similarity functions</strong>&nbsp;to align datasets with high accuracy and strong coverage.&nbsp;</li> <li>Define and track&nbsp;<strong>data quality indicators</strong>, such as overlap metrics, match precision/recall, duplicate rates, and completeness.&nbsp;</li> <li>Prepare training-ready datasets in formats such as&nbsp;<strong>TFRecords</strong>, and structure data to meet ML research requirements.&nbsp;</li> <li>Develop processing components using&nbsp;<strong>Dataflow (Beam)</strong> and manage large analytical workloads in <strong>BigQuery</strong>.&nbsp;</li> <li>Leverage frameworks like&nbsp;<strong>Ray</strong>&nbsp;to accelerate large-scale experiments, feature extraction, and research-oriented data preparation.&nbsp;</li> <li>Collaborate with ML researchers to anticipate downstream requirements and evolve linkage strategies as new sources and use cases emerge.&nbsp;</li> </ul> <h2>What We&#