Jobs and Careers
QU

Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)

qualcomm
San Diego, United Statesfull_timeVerifiedPosted 3 Jun 2026
💰 $237,600/yr($158,400/yr$237,600/yr)

About the role


Company:

Qualcomm Technologies, Inc.

Job Area:

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary:

THIS IS A FULL-TIME ONSITE ROLE REQUIRING 5 DAYS A WEEK IN OFFICE AT QUALCOMM’S SAN DIEGO LOCATION

As a leading technology innovator, Qualcomm pushes the boundaries of what’s possible to enable next-generation experiences and drive digital transformation, creating a smarter, connected future for all.

As a Staff/Sr. Staff Software Engineer in the Qualcomm AI Stack SDK Software team, you will design, develop, and deliver advanced AI/ML software solutions for Generative AI inference on Snapdragon platforms. This role focuses on model optimization, quantization, graph transformations, and runtime execution for modern AI architectures including LLMs, LVMs, and LMMs.

You will work at the intersection of machine learning algorithms, inference optimization, graph lowering, and systems software, contributing directly to the Qualcomm AI Stack SDK (QAIRT), and associated tools, including delegates support for ONNX Runtime, Executorch and TFLite/LiteRT frameworks. You will collaborate with amazing engineers from different teams across multiple locations like ML Research, AI accelerator HW/SW teams, Product Management, Program Management, and QA to drive features from concept to production.

This role requires strong technical ownership, the ability to work independently, and the capability to drive features end-to-end while mentoring junior engineers.

Minimum Qualifications:

• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.

Responsibilities:

• Convert, optimize, and deploy AI models from PyTorch and ONNX frameworks for efficient inference on Snapdragon platforms.

• Design and implement graph transformations, graph lowering, and optimization techniques within AI runtime environments such as ONNX Runtime, ExecuTorch and Qualcomm AI Stack SDK.

• Apply knowledge of quantization and performance optimization to improve latency, throughput, memory usage, and power efficiency.

• Work at the forefront of Generative AI, understanding advanced algorithms such as attention mechanisms, Mixture-of-Experts (MoE), Low Rank Adapter (LoRA) and emerging inference optimization techniques (e.g., Speculative Decoding etc.).

• Collaborate with ML Research teams to prototype and productize new features and techniques into SDK solutions.

• Debug complex issues across models, runtime, OS, compiler, and hardware layers, working closely with QA and customer teams.

• Design, implement, and deliver new features and enhancements to the Qualcomm AI Stack SDK.

• Participate in design reviews and code reviews, ensuring software quality and maintainability.

• Mentor junior engineers helping them prioritize work, and drive execution across multiple initiatives.

Minimum Qualification (Must Have)

• Bachelor’s degree in computer science, computer engineering, or a related field and 6+ years (Staff) / 8+ years (Sr. Staff) of experience in software design, development, and delivery.

• OR

• Master’s degree/PhD in computer science, computer engineering, or a related field and 5+ years (Staff) / 7+ years (Sr. Staff) of experience in software design, development and delivery.

• 3+ years of hands-on experience in AI/ML software development, with a focus on inference or model optimization.

• Strong understanding of AI/ML fundamentals, including deep learning and inference pipelines.

• Deep understanding of transformer architectures, attention mechanisms, and performance tradeoffs.

• Proficiency in Python and C/C++ for production-quality software development.

• Experience working with PyTorch and ONNX models and tooling.

• Debugging skill of complex issues, perform root cause analysis, and ensure high system reliability

• Ability to work independently, collaborate across teams, and drive complex features end-to-end.

Preferred Qualifications

• Working knowledge of graph theory, graph optimizations, and compiler-

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

qualcomm

View company profile →