LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer

Qualcomm

San Diego (CA)

On-site

USD 158,400 - 237,600

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive annual discretionary bonus program
Potential RSU grants
Comprehensive benefits package

Job summary

A leading technology firm in San Diego seeks an LLM Serving Engineer to develop scalable AI solutions. This role involves building LLM inference platforms and collaborating with teams to drive innovations in machine learning. Responsibilities include optimizing deep learning workloads and utilizing advanced techniques for efficient serving. Candidates should have strong experience with LLM packages, a solid foundation in computer science, and experience in Python development. Competitive salary and benefits are offered.

Qualifications

  • Hands-on experience with Triton-Inference Server and similar packages.
  • Strong experience in developing language models, especially using PyTorch.
  • Excellent understanding of algorithms and parallel programming.

Responsibilities

  • Build a scalable LLM inference platform using advanced techniques.
  • Contribute to development of LLM Serving packages.
  • Drive efficient serving with load balancing and routing.

Skills

Experience with LLM serving packages
Deep understanding of foundational LLMs
Experience in developing language models using PyTorch
Computer science fundamentals
Understanding of computer architecture and ML accelerators
Python development skills
Experience in optimizing deep learning workloads
Problem-solving skills
Excellent communication skills

Education

Bachelor’s degree in relevant field
Master’s degree in relevant field
PhD in relevant field

Job description

Company

Qualcomm Technologies, Inc.

Job Area

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary

LLM Serving Engineer (Cloud AI Engineering)

Qualcomm is utilizing its traditional strengths in digital wireless technologies to play a central role in the evolution of Cloud AI. We are investing in several supporting technologies including Deep Learning. The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration.

We are hiring LLM Serving Engineers at multiple levels to join our dynamic, collaborative team. This role spans the full product lifecycle—from cutting-edge research and development to commercial deployment—and demands strategic thinking, strong execution, and excellent communication skills.

Role Activities
  • Building a scalable LLM inference platform using inference techniques (e.g. disaggregated serving and KV-Cache management, advanced parallelism, speculative algorithms, model optimization, specialized kernels).
  • Contribute to the development of LLM Serving packages (e.g. vLLM, SGLang, TGI, Triton-Inference Server, Dynamo, LLM-d).
  • Work closely with customers to drive solutions by collaborating with internal compiler, firmware and platform teams.
  • Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics to identify new optimization opportunities.
  • Drive efficient serving through autoscaling, load balancing and routing.
  • Engage with open-source serving communities to evolve the framework.
Qualifications
  • Hands-on experience in one or more of the following LLM serving/orchestration packages (Triton-Inference Server, vLLM, SGLang, Ollama, llm-d, KServe, LMCache, MoonCake).
  • Deep understanding of foundational LLMs, VLMs, SLMs, transformer-based architectures.
  • Strong experience in developing language models using PyTorch.
  • Strong computer science fundamentals - algorithms, data structures, parallel and distributed programming.
  • Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
  • Strong Python development skills for large-scale projects and a passion for software engineering.
  • Experience in analyzing, profiling, and optimizing deep learning workloads.
  • Proactive learning about the latest inference optimization techniques.
  • Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.
  • MS or BS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • OR Master’s degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • OR PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Additional Information

Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, Qualcomm is committed to providing an accessible process. You may email disability-accomodations@qualcomm.com or call Qualcomm’s toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities.

EEO Employer: Qualcomm is an equal opportunity employer; all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or any other protected classification.

Pay range and Other Compensation & Benefits

$158,400.00 - $237,600.00

The above pay scale reflects the broad, minimum to maximum, pay scale for this job code for the location for which it has been posted. Salary is only one component of total compensation at Qualcomm. We offer a competitive annual discretionary bonus program and potential RSU grants. Our benefits package supports employees at work, at home, and at play. Your recruiter can discuss details about Qualcomm benefits.

If you would like more information about this role, please contact Qualcomm Careers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer
LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
Competitive salary
Annual discretionary bonus
RSU grants opportunity
Senior Engineer - Machine Learning
Senior Engineer - Machine Learning

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive annual discretionary bonus program
Opportunity for annual RSU grants
Comprehensive benefits package
Machine Learning Engineer, Staff (Model Optimization)
Machine Learning Engineer, Staff (Model Optimization)

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Senior Systems Engineer, Data Center AI
Senior Systems Engineer, Data Center AI

Qualcomm • San Diego (CA)

On-site
USD 111,000 - 167,000
AI Researcher, On-Device LLM Efficiency
AI Researcher, On-Device LLM Efficiency

Qualcomm • San Diego (CA)

On-site
USD 138,000 - 209,000
Competitive discretionary bonus
Annual RSU grants
Benefits package
Machine Learning Engineer, Staff (Model Optimization) San Diego, California, United States of America Machine Learning Engineering
Machine Learning Engineer, Staff (Model Optimization) San Diego, California, United States of America Machine Learning Engineering

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Machine Learning Engineer - College Graduate
Machine Learning Engineer - College Graduate

Latitude • San Diego (CA), Northern (KY)

Hybrid
USD 123,000 - 184,000
Machine Learning Engineer - Generative AI
Machine Learning Engineer - Generative AI

Latitude • San Diego (CA), Northern (KY)

Hybrid
USD 104,000 - 156,000
Sr. Engineer, AI Platforms and Solutions
Sr. Engineer, AI Platforms and Solutions

Qualcomm • San Diego (CA)

On-site
USD 111,000 - 167,000
Sr. Engineer, AI Platforms and Solutions
Sr. Engineer, AI Platforms and Solutions

JobCubby • San Diego (CA)

On-site
USD 111,000 - 167,000