Jobs and Careers
CA

Senior Data Engineer - Senior Spark Data Engineer

Capgemini
United Statesfull_timeVerifiedPosted 29 May 2025

About the role

Senior Data Engineer - Senior Spark Data Engineer-079918

Description

 

About the job you’re considering

We are seeking an experienced Senior Data Engineer with strong expertise in Apache Spark to help build and manage scalable, secure, and efficient data platforms. This role will be instrumental in designing data architectures and pipelines that support both advanced analytics and governed data access across the organization. You will work with cross-functional teams to enable data discovery, lineage, and compliance while delivering high-performance data processing systems.

Your role

  • Design, build, and optimize scalable data pipelines using Apache Spark.
  • Manage and govern data access and metadata using AWS Data Zone.
  • Implement and enforce data access controls, lineage tracking, and data classification.
  • Integrate data across cloud platforms and on-prem systems into unified data lakes and warehouses.
  • Partner with data analysts, scientists, and product teams to deliver clean, reliable, and well-governed datasets.
  • Develop and automate ingestion, transformation, and quality validation workflows.
  • Ensure data compliance and security policies are implemented consistently across the platform.
  • Contribute to architecture and governance strategy for enterprise-scale data platforms.
  • Support performance tuning, troubleshooting, and monitoring of Spark jobs and data pipelines.
  • Mentor junior engineers and support team development through code reviews and documentation.

Your skills and experience

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in data engineering with production-level data systems.
  • Expert in Apache Spark (PySpark or Scala), including performance tuning and optimization.
  • Strong experience with Datazone for data governance and access management.
  • Proficient in SQL and modern data architecture concepts (e.g., lakehouse, Delta Lake).
  • Hands-on experience with cloud platforms (AWS, Azure, or GCP), especially in data services (e.g., S3, ADLS, Redshift, Synapse).
  • Experience with orchestration tools like Airflow, Athena, or similar.
  • Building the Frame work to load data for the data Ingestion for the sources files & RDBMS
  • Designing and developing data layers using Framework configuration by using the ABCR Meta data.
  • Using PySpark/Scala to load data, created schema, processed data and sent to kafka.
  • Optimization of Spark Jobs using Pyspark.
  • Performing data processing such as aggregation, joins, filter as per business rule.
  • Strong knowledge of data governance, lineage, access control, and compliance frameworks.
  • Familiarity with DevOps and infrastructure-as-code tools (e.g., Terraform, Git, CI/CD pipelines).

Life at Capgemini

Capgemini supports all aspects of your well-being throughout the changing stages of your life and career. For eligible employees, we offer:

  • Flexible work
  • Healthcare including dental, vision, mental health, and well-being programs
  • Financial well-being programs such as 401(k) and Employee Share Ownership Plan
  • Paid time off and paid holidays
  • Paid parental leave
  • Family building benefits like adoption assistance, surrogacy, and cryopreservation
  • Social well-being benefits like subsidized back-up child/elder care and tutoring
  • Mentoring, coaching and learning programs
  • Employee Resource Groups
  • Disaster Relief

<

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Capgemini

View company profile →