Jobs and Careers
AD

Principal Scientist - Data Pipeline Engineer

Adobe
San Jose, United StatesRemotefull_timeVerifiedPosted 17 Jul 2026
💰 $388,000/yr($206,300/yr$388,000/yr)

About the role

ABOUT THE ROLE 

We’re looking for a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Adobe Firefly’s multimodal foundation models (image, video, audio). In this role, you’ll sit at the intersection of data engineering and applied ML building distributed, GPU-accelerated systems that turn billions of raw assets into training-ready data at scale. 

Your work will directly determine how fast and how well Adobe models can learn directly impacted by the throughput and reliability of our data pipelines, and the quality of data that reaches training. This is a senior individual contributor role with broad technical influence across data, infrastructure, and modeling teams. 

WHAT YOU’LL DO 

OPTIMIZE DATA PROCESSING PIPELINES AT SCALE 

  • Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training-ready data 

  • Scale up inference throughput across the pipeline (batching, parallelism, hardware utilization) to turn raw collected data into training data faster and more cheaply 

  • Identify and eliminate bottlenecks across ingestion, processing, and delivery, from storage and I/O to compute scheduling 

ARCHITECT SCALABLE DATA INFRASTRUCTURE 

  • Design systems that reliably store, index, and serve billions of data points, each requiring substantial processing spanning large-scale databases, distributed storage, and high-throughput compute 

  • Apply deep expertise in distributed systems and frameworks such as Ray (or equivalent) to orchestrate large-scale, GPU/CPU-heavy data workloads 

  • Own architecture decisions including database and storage choices, job scheduling, GPU cluster utilization that let the platform scale alongside data and model growth 

DRIVE DATA CURATION FOR MODEL TRAINING 

  • Bring a strong ML background, especially inference optimization for VLMs and LLMs and data curation for training 

  • Partner closely with modeling teams to understand what data improves training outcomes, and translate that into pipeline and curation requirements 

  • Operate as a hands-on technical leader who bridges data engineering and applied ML 

WHAT YOU NEED TO SUCCEED 

  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including work at large scale (billions of records or assets) 

  • Strong software engineering background, with hands-on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large-scale 

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Adobe

View company profile →