Lead AWS Data Engineer
Publicis GroupeAbout the role
Job Description
About you:
As a lead data engineer, you will design and maintain data platform road maps and data structures that support business and technology objectives. Naturally inquisitive and open to the deep exploration of underlying data, finding actionable insights, and working with functional competencies to drive identified actions. You also enjoy working both freely and as part of a team and have the confidence to influence and communicate with stakeholders at all levels, and to work in a fast-paced complex environment with conflicting priorities.
About the role:
Reporting into the delivery leader, you will deliver consumable, contemporary, and immediate data content to support and drive business decisions. The key focus of the role is to deliver a custom solution to support various business critical requirements. You will be involved in all aspects of data engineering from delivery planning, estimating and analysis, all the way through to data architecture and pipeline design, delivery, and production implementation. From the beginning you will be involved in the design and implementation of complex data solutions ranging from batch to streaming and event-driven architectures, across cloud, on-premise, and hybrid client technology landscapes.
Brief Description of Role:
We are looking for 5+ years of experience in data engineering in a customer or business facing capacity and experience in the following:
- Ability to understand and articulate requirements to technical and non-technical audiences
- Stakeholder management and communication skills, including prioritizing, problem solving and interpersonal relationship building
- Strong experience in SDLC delivery, including waterfall, hybrid, and Agile methodologies
- Experience of implementing and delivering data solutions and pipelines on AWS Cloud Platform.
- A strong understanding of data modelling, data structures, databases, and ETL processes
- An in-depth understanding of large-scale data sets, including both structured and unstructured data
- Knowledge and experience of delivering CI/CD and DevOps capabilities in a data environment
- Develop new inbound data source ingestions required within the multi-tiered data platform to support analytics and marketing automation solutions
- Supports data pipelines – Builds the required dimensions, rules, segments, and aggregates
- Support all database operations: performance monitoring, pipeline ingestion, maintenance, etc.
- Monitor platform health - data loads, extracts, failures, performance tuning
- Create/modify data structures/pipelines
- Leveraging capabilities of Databricks Lakehouse functionality as needed to build Common/Conformed layers within the data lake
- Application of software best practices into consideration. E.g. infrastructure cost, security, scalability, feature extensibility, ease of maintenance, etc.
- Develop, document, and test software and environment setup to ensure that the outcome meets the needs of end-users and achieves business goals
Qualifications
The following skills are required:
- Tech Stack: AWS pipeline, Glue, Databricks, Python, SQL, Spark, etc.
- Building the Data Lake using AWS technologies like S3, EKS, ECS, AWS Glue, AWS KMS, EMR
- Extensive experience in ETL and audience segmentation
- Developing sustainable, scalable, and adaptable data pipelines
- Attention to detail in design, documentation, and test coverage of delivered tasks
- Strong written and verbal communication skills, team player
- Strong analytical and leadership skills
- In addition, the candidate should have strong business acumen, interpersonal skills, and communication skills, yet also be able to work independently.
- At least 4 years of experience with designing and developing Data Pipelines for Data Ingestion or Transformation using AWS technologies
- At least 3 years of experience in the following Big Data frameworks: File Format (Parquet, etc.), Resource Management, Distributed Processing
- At least 4 years of experience developing applications with Monitoring, Build Tools, Version Control, Unit Test, TDD, Change Management to support DevOps
- At least 4 years of experience with Spark programming (PySpark or Scala)
- At least 4 years of experience with Databricks implementations
- Familiarity with the concepts of “delta lake” and “lakehouse” technologies
- At least 3 years of experience working with Streaming using Spark or Flink or Kafka
The following skills are nice to have, and expertise is not required:
- Adobe (Campaign, Audience Manager, Analytics)
- MLFlow
- Microsoft Power BI
- SAP Business Objects
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s