Data Engineer Principal_3003
AllianzAbout the role
We are looking for an experienced Principal Data Engineer to work hands-on on the re-engineering of an existing enterprise data platform built on Azure Synapse Analytics. The role requires strong technical depth to audit, understand, and validate a complex end-to-end data architecture spanning source ingestion through to consumption — and to help deliver the migration of validated workloads to Databricks. This is hands on role and demands the ability to reverse-engineer existing implementations, assess their correctness, and build to an agreed migration strategy.
Key Responsibilities
- Contribute to the technical assessment and re-engineering of an existing enterprise data platform, spanning all layers from source ingestion through to data consumption
- Reverse-engineer, document, and validate existing pipeline logic, data models, transformation frameworks, and data governance controls
- Identify gaps, defects, and technical debt across the platform and remediate where implementations are incorrect or sub-optimal
- Ensure correctness of data processing patterns including change data capture, slowly changing dimensions, deduplication, and business reconciliation
- Implement target-state designs aligned to modern lakehouse principles, ensuring feature parity and business logic fidelity during transitions
- Support platform evolution initiatives, including parallel-run phases where multiple implementations operate simultaneously, validating output consistency before cutover
- Execute migration of existing workloads to modern data platforms, preserving existing governance and control framework semantics
- Re-implement ingestion, transformation, and orchestration pipelines on target platforms, maintaining audit, quality, and reconciliation standards
- Collaborate with business, data governance, and architecture stakeholders to validate embedded business rules and data quality requirements
- Mentor less experienced engineers, review code and designs, and support decommission planning for legacy components
Core Technical Skills
- Azure Synapse & Data Platform Mandatory hands-on expertise with:
- Azure Synapse Analytics (Pipelines, Spark Pool, Dedicated SQL Pool)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Delta Lake on Azure (Synapse Lakehouse patterns)
- Oracle Golden Gate Replication for real-time source integration
- Azure Analysis Services and Power BI consumption layer patterns
- Deep understanding of medallion architecture: Raw / Harmonized / Conformed / Consumption layers
- Strong knowledge of SCD Type 0/1/2, CDC patterns, soft/hard delete, and retroactive change processing
- Experience with Synapse SQL Pool — stored procedures, control tables, and data quality validation patterns
- Experience with audit, balance, and control frameworks — parameterized, modular pipeline governance at enterprise scale
- Familiarity with config-driven and automation-first pipeline patterns (YAML, PySpark, SQL-driven generation from mapping documents)
Databricks & Lakehouse
- Hands-on experience with Azure Databricks (Delta Live Tables, Unity Catalog preferred)
- Strong Apache Spark skills (PySpark / Spark SQL)
- Experience migrating workloads from legacy data warehouse or Synapse environments to a Databricks Lakehouse
- Ability to re-implement governance and control frameworks natively in Databricks (audit logging, reconciliation, DQ checks)
- Experience with Delta Lake features: MERGE, CDC, time travel, schema enforcement
Data Engineering & Development
- Strong Python and SQL programming skills
- Experience with ETL/ELT at scale: denormalization, surrogate keys, directory tables, curated data models
- Experience integrating complex data sources: Oracle DB, SQL Server, Azure SQL DB, file systems, Salesforce, APIs
- Strong data modelling skills: relational, dimensional, and lakehouse-oriented
DevOps & Automation
- CI/CD pipelines for data engineering (Azure DevOps / GitHub Actions)
- Infrastructure as Code (Terraform or ARM)
- Containerization (Docker)
- Experience with automated testing frameworks for data pipelines (unit testing, reconciliation-based validation)
Nice to Have
- Experience with Unity Catalog for data governance and lineage
- Familiarity with Azure Purview for data cataloguing and governance
- Exposure to real-time and streaming pipelines (Event Hub / Kafka / Kinesis)
- Experience with GenAI or ML p
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s