Jobs and Careers
3M

Principal Data Engineer

3M
United Statesfull_timeVerifiedPosted 4 Dec 2024
💰 $195,903/yr($160,284/yr$195,903/yr)

About the role

Job Description:

Principal Data Engineer

Collaborate with Innovative 3Mers Around the World

Choosing where to start and grow your career has a major impact on your professional and personal life, so it’s equally important you know that the company that you choose to work at, and its leaders, will support and guide you. With a diversity of people, global locations, technologies and products, 3M is a place where you can collaborate with other curious, creative 3Mers.

This position provides an opportunity to transition from other private, public, government or military experience to a 3M career.

The Impact You’ll Make in this Role

3M is seeking a Principal Data Engineer to join the Corporate Research Systems Lab (CRSL) to develop scalable Data Systems. As part of an agile team, you will enable applications in diverse markets including energy, manufacturing, personal safety, transportation, electronics, and consumer.

As a Principal Data Engineer, you will have the opportunity to design and support an Enterprise Data Mesh to empower informatics and digital technologies for users across the globe:

  • Architect, design, and build scalable, efficient, and fault-tolerant data operations.
  • Collaborate with senior leadership, analysts, engineers, and scientists to implement new mesh domain nodes and data initiatives.
  • Drive technical architecture for accelerated solution designs, including data integration, modeling, governance, and applications.
  • Explore and recommend new tools and technologies to optimize the data platform.
  • Improve and implement data engineering and analytics engineering best practices.
  • Collaborate with data engineering and domain nodes teams to design physical data models and mappings.
  • Work with scientists and informaticians to develop advanced digital solutions and promote digital transformation and technologies.
  • Perform code reviews, manage code performance improvements, and enforce code maintainability standards.
  • Develop and maintain scalable data pipelines for ingesting, transforming, and distributing data streams.
  • Advise and mentor 3M businesses, data scientists, and data consumers on data standards, pipeline development, and data consumption.
  • Provide technical guidance and mentorship, ensure adherence to best practices, and maintain high software quality through rigorous testing and code reviews.
  • Guide project planning and execution, manage timelines and resources, and facilitate effective communication between team members and stakeholders.
  • Foster a positive team environment, assist in recruitment, and provide training opportunities to address skill gaps.

Your Skills and Expertise  

To set you up for success in this role from day one, 3M requires (at a minimum) the following qualifications:

  • Bachelor’s degree or higher in Computer Science from an accredited university. 
  • Ten (10) years of professional experience in data management, data engineering, data governance, and data warehouse/lakehouse design and development with proficiency. across SQL and NoSQL data management systems and having comfort working with structured and unstructured data and analyses.
  • Five (5) years of extensive experience and proficiency with Python, Apache Spark, PySpark, and Databricks
  • Three (3) of hands-on experience in Python to extract data from APIs, build data pipelines.

Additional qualifications that could help you succeed even further in this role include:

  • Exposure to data and data types in the Materials science, chemistry, computational chemistry, physics space.
  • Proficiency in developing or architecting modern distributed cloud architecture and workloads (AWS, Databricks preferred). Familiarity with data mesh style architecture design principles.
  • Proficiency in building data pipelines to integrate business applications and procedures.
  • Solid understanding preferred of advanced Databricks concepts like Delta Lake, MLFlow, Advanced Notebook Features, Custom Libraries and Workflows, Unity Catalog, etc.
  • Experience with AWS cloud computing services and infrastructure developing data lakes and data pipelines leveraging multiple technologies such as AWS S3, AWS Glue, Elastic MapReduce, etc. and awareness of considerations for building scalable, distributed computational systems on Spark.
  • Experience with stream-processing systems: Amazon Kinesis, Spark, Storm, Kafka, etc.
  • Data quality and validation principles experience, security principles data encryption, access control, authentication & authorization.
  • Deep experience in definition and implementation of feature engineering.
  • Experience with Docker contain

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

3M

View company profile →