Lead Machine Learning Engineer, Performance and Scalability, Generative AI
AdobeAbout the role
Our Company
Changing the world through digital experiences is what Adobe’s all about. We give everyone—from emerging artists to global brands—everything they need to design and deliver exceptional digital experiences! We’re passionate about empowering people to create beautiful and powerful images, videos, and apps, and transform how companies interact with customers across every screen.
We’re on a mission to hire the very best and are committed to creating exceptional employee experiences where everyone is respected and has access to equal opportunity. We realize that new ideas can come from everywhere in the organization, and we know the next big idea could be yours!
About the Role
Adobe Firefly is seeking a Lead Engineer to focus on Performance and Scalability for our Generative AI systems, powering flagship products like Photoshop, Illustrator, Express, and firefly.adobe.com. In this senior role, you will be responsible for optimizing high-performance, scalable AI pipelines, supporting millions of users worldwide.
You will work closely with machine learning researchers, infrastructure engineers, and applied scientists to ensure that generative AI models are efficiently deployed, scaled, and monitored, without directly implementing model training, quantization, or tensor parallelism.
Responsibilities
Architect and optimize ML pipelines to support scalable inference and model deployment on cloud-based GPU infrastructure (e.g., AWS P5 instances).
Develop and maintain high-throughput serving pipelines for generative AI models, ensuring low-latency, high-performance execution.
Enable model serving optimizations by designing systems that support tensor parallelism, quantization, distillation, and caching, in collaboration with ML research teams.
Develop automated monitoring and profiling tools to track system efficiency, detect performance regressions, and optimize resource utilization.
Optimize GPU resource allocation and orchestration across cloud-based ML workloads.
Integrate scalable load testing frameworks to validate model inference performance under high-traffic conditions.
Collaborate with infrastructure and applied ML teams to transition models from experimentation to production-ready, cloud-optimized deployments.
Establish standard methodologies for scaling and cloud-native ML architectures, ensuring efficient deployment across multi-region cloud environments.
Qualifications
8+ years of proven track record in building high-performance ML infrastructure and scalable AI systems.
MS, or PHD in computer science or related field.
Strong programming skills in Python and C++, with expertise in building ML pipelines and model deployment in
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s