Staff AI Model Optimization Architect for LLMs & Multimodal

Qualcomm

Austin (TX)

On-site

USD 158,000 - 238,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Qualcomm Technologies, Inc. seeks a Staff Engineer to lead end-to-end AI model optimization for LLMs, VLMs, and diffusion on Qualcomm accelerators. You will transform PyTorch models, manage KVcache behavior, and drive deployment with PyTorch, ONNX, and torch.compile across multi-core systems.

You will collaborate with compiler and performance teams to craft lowering strategies and scalable tooling, ensuring high throughput with low latency and production-grade reliability.

Qualifications

  • Bachelor's degree in CS/Engineering with 4+ years of related work experience.
  • Master's degree with 3+ years of related work experience.
  • PhD in CS/Engineering or related field with 2+ years of related work experience.
  • Experience in hardware/software systems and ML accelerators preferred.

Responsibilities

  • Architect and deliver model optimization strategies for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile.
  • Design fusion kernels via DSLs (e.g., Triton) and perform kernel fusion.
  • Collaborate with compiler, performance, and accuracy teams on lowering strategies.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency.
  • Own transformer optimizations including KVcache management and decoding behavior.
  • Support distributed inference strategies across multi-core/multi-device systems.
  • Create reusable patterns and tooling to scale model optimizations to new hardware.
  • Debug complex performance issues to deliver production-ready solutions.

Skills

PyTorch
Python
Transformer architectures
KVcache management
Distributed systems
Model optimization
Research leadership

Education

Bachelor's degree in CS/Engineering
Master's degree in CS/Engineering
PhD in CS/Engineering

Tools

TorchDynamo
Torch.compile
ONNX
Triton

Job description

Qualcomm Technologies, Inc. seeks a Staff Engineer to lead end-to-end AI model optimization for LLMs, VLMs, and diffusion on Qualcomm accelerators. You will transform PyTorch models, manage KVcache behavior, and drive deployment with PyTorch, ONNX, and torch.compile across multi-core systems.

You will collaborate with compiler and performance teams to craft lowering strategies and scalable tooling, ensuring high throughput with low latency and production-grade reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Model Optimization Architect (LLM & Multimodal)
Staff AI Model Optimization Architect (LLM & Multimodal)

QUALCOMM, Inc. • Austin (TX)

On-site
USD 162,000 - 243,000
AI Model Optimization Architect
AI Model Optimization Architect

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
AI Model Optimization Architect
AI Model Optimization Architect

QUALCOMM, Inc. • Austin (TX)

On-site
USD 162,000 - 243,000
AI Hardware Performance Modeling Engineer for LLMs
AI Hardware Performance Modeling Engineer for LLMs

Qualcomm • San Diego (CA)

On-site
USD 162,000 - 243,000
AI Hardware Performance Modeling for LLMs
AI Hardware Performance Modeling for LLMs

QUALCOMM, Inc. • San Diego (CA)

On-site
USD 162,000 - 243,000
Annual bonus
RSU grants
Comprehensive benefits
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff

Qualcomm • San Diego (CA)

On-site
USD 178,400 - 267,600
Competitive annual discretionary bonus
RSU grants
Highly competitive benefits package
Senior/Staff AI Performance Engineer: Inference Optimization
Senior/Staff AI Performance Engineer: Inference Optimization

Qualcomm • San Diego (CA)

On-site
USD 178,400 - 267,600
Competitive annual discretionary bonus
RSU grants
Highly competitive benefits package
Staff Machine Learning Engineer – AI/ML Compiler
Staff Machine Learning Engineer – AI/ML Compiler

Qualcomm • Santa Clara (CA)

On-site
USD 160,500 - 240,700
Performance Modeling Engineer for AI Accelerators and SoC
Performance Modeling Engineer for AI Accelerators and SoC

Qualcomm Technologies • Austin (TX)

On-site
USD 192,000 - 288,000
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000