Jobs and Careers
QU

Staff Machine Learning Engineer – Model Optimization & Quantization

qualcomm
San Diego, United Statesfull_timeVerifiedPosted 11 Mar 2026
💰 $237,600/yr($158,400/yr$237,600/yr)

About the role


Company:

Qualcomm Technologies, Inc.

Job Area:

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary:

About the Role 

Join the Qualcomm AI Hub team and help developers integrate machine learning into their products and experiences: https://aihub.qualcomm.com/. 

In this role you will develop tools to help developers optimize and deploy machine learning models on edge and mobile hardware. AIMET is Qualcomm's open-source library for state-of-the-art model quantization, and compression techniques. You will develop and support cutting-edge model optimization workflows — pushing the boundary of what's possible on resource-constrained hardware. Applications range from quantizing large language models (LLMs) and generative AI models to compressing latency-critical vision, audio, and multimodal networks for deployment on Qualcomm Snapdragon and other edge SoCs. 

For this role we are seeking a talented and motivated Staff Software Engineer with expertise in the optimizing and deploying ML models – especially for edge devices 

 

What You'll Do 

  • Design, develop, and maintain quantization algorithms and compression pipelines within the AIMET framework (PTQ, QAT, mixed-precision, AdaScale etc.) 

  • Implement advanced quantization techniques including weight-only quantization, activation quantization, KV-cache quantization, and sub-4-bit quantization for LLMs and generative AI models 

  • Build tooling to analyze, profile, and debug model accuracy degradation caused by quantization 

  • Integrate AIMET workflows with popular ML frameworks — PyTorch and ONNX 

  • Develop APIs and developer-facing tooling to make AIMET accessible and easy to use for external customers and design partners 

  • Integrate AIMET in AI Hub Workbench Quantize job to enable Quantization at large scale. 

  • Own end-to-end quantization and optimization of models published on Qualcomm AI Hub, ensuring they meet accuracy, latency, and power targets on Qualcomm hardware 

  • Quantize and validate a broad range of model families — vision transformers, LLMs, diffusion models, speech, and multimodal architectures — for deployment via AI Hub 

  • Develop and maintain automated quantization pipelines and evaluation harnesses to scale model onboarding across AI Hub's growing model catalog 

 

Minimum Qualifications:

• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
PhD in Computer Science, Engineering, Information Systems, or related field and 2+ yea

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

qualcomm

View company profile →