Senior Data Engineer - Consumer Data AI / ML
YahooAbout the role
Summary:
We are looking for a Senior Data Engineer to design, build and optimize the scalable data pipelines that power our advanced analytics and machine learning solutions. In this role, you will collaborate closely with data scientists, software engineers, and business stakeholders to prepare and transform large datasets, support end-to-end model development and deployment, and ensure robust, efficient, and secure data flows. You will leverage your expertise in cloud platforms, big data tools, and machine learning frameworks to drive innovation and deliver actionable insights that advance our organization’s AI initiatives and business objectives.
Responsibilities:
Design, build, and maintain scalable data pipelines and ETL processes to support machine learning and AI initiatives on Google Cloud Platform (GCP).
Implement and optimize data storage solutions using GCP services such as BigQuery, Cloud Storage, and Dataflow.
Ensure data quality, integrity, and security throughout its entire lifecycle.
Collaborate with data scientists, analysts, and business stakeholders to understand data requirements and deliver actionable insights.
Troubleshoot and optimize the health of cloud-based data infrastructure to ensure reliability.
Automate manual processes and repetitive tasks to improve efficiency and reduce errors.
Apply data governance and compliance best practices to protect sensitive information and meet regulatory standards.
Document solutions, processes, and architectural decisions to facilitate knowledge sharing and maintainability.
Qualifications:
BS or MS in Computer Science or a related field, or equivalent experience.
5+ years of experience in Data Engineering, with a heavy emphasis on backend data systems and pipeline design.
3+ years hands-on experience with Google Cloud Platform ecosystem (BigQuery, Dataproc, Composer, Dataflow, Data Catalog, Observability) or AWS equivalent.
Proven ability to design, build, and maintain data pipelines that support machine learning and AI model development, training, and deployment.
High proficiency in Java (preferred) or Python, along with proficiency in SQL for database operations.
Familiarity with data security, compliance, and governance best practices.
Strong problem-solving skills, attention to detail, and ability to work collaboratively with cross-functional teams.
Excellent communication skills and ability to tell insightful stories using data and also manage communication within internal teams and stakeholders.
Bonus (Not Required): Previous exposure to or interest in building pipelines for Machine Learning (ML) model training and deployment.
The material job duties and responsibilities of this role include those listed above as well as adhering to Yahoo policies; exercising sound judgment; working effectively, safely and inclusively with others; exhibiting trustworthiness and meeting expectations; and safeguarding business operations and brand integrity.
At Yahoo, we offer flexible hybrid work options that our employees love! While most roles don’t require regular office attendance, you may occasionally be asked to attend in-person events or team sessions. You’ll always get notice to make arrangements. Your recruiter will let you know if a specific job requires regular attendance at a Yahoo office or facility. If you have any questions about how this applies to the role, just ask the recruiter!
Yahoo is proud to be an equal opportunity workplace. All qualified applicants will receive consideration for employment without regard to, and will not be discriminated against based on age, race, gender, color, religion, national origin, sexual orientation, gender identity, veteran status, disability or any other protected category. Yahoo will consider
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s