Jobs and Careers
QU

LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer

qualcomm
San Diego, United Statesfull_timeVerifiedPosted 9 Apr 2026
💰 $237,600/yr($158,400/yr$237,600/yr)

About the role


Company:

Qualcomm Technologies, Inc.

Job Area:

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary:

LLM Serving Engineer (Cloud AI Engineering) 

 

Qualcomm is utilizing its traditional strengths in digital wireless technologies to play a central role in the evolution of Cloud AI. We are investing in several supporting technologies including Deep Learning. The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration.  

 

We are hiring LLM Serving Engineers at multiple levels to join our dynamic, collaborative team. This role spans the full product lifecycle—from cutting-edge research and development to commercial deployment—and demands strategic thinking, strong execution, and excellent communication skills. 

 

This role involves the following activities: 

  • Building a scalable LLM inference platform using inference techniques (e.g. disaggregated serving and KV-Cache management, advanced parallelism, speculative algorithms, model optimization, specialized kernels). 

  • Contribute to the development of LLM Serving packages (e.g. vLLM, SGLang, TGI, Triton-Inference server, Dynamo, LLM-d). 

  • Work closely with customers to drive solutions by collaborating with internal compiler, firmware and platform teams. 

  • Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics to identify new optimization opportunities. 

  • Drive efficient serving through smart autoscaling, load balancing and routing. 

  • Engage with open-source serving communities to evolve the framework. 

 

Candidates for this position will demonstrate the following: 

  • Hands-on experience in one or more of the following LLM serving/Orchestration packages (Triton-Inference Server, vLLM, SGLang, Ollama, llm-d, KServe, LMCache, MoonCake) 

  • Deep understanding of foundational LLMs, VLMs, SLMs, transformer-based architectures. 

  • Strong experience in developing language models using PyTorch. 

  • Strong computer science fundamentals - algorithms, data struc

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

qualcomm

View company profile →