Staff AI Model Optimization Architect (LLM & Multimodal)

QUALCOMM, Inc.

Austin (TX)

On-site

USD 162,000 - 243,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Qualcomm Technologies, Inc. in Austin, TX, is seeking a Staff Engineer – AI Model Optimization Architect to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm accelerators.

You will collaborate with compiler, performance, and accuracy teams to translate models into accelerator-efficient execution, balancing throughput, latency, memory, and quality across Day0 to production deployment.

Qualifications

  • Expert level PyTorch and inference-focused model optimization.
  • Experience with torch.compile / TorchDynamo workflows.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
  • Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
  • Strong foundation in computer architecture, ML accelerators, and distributed systems.

Responsibilities

  • Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
  • Design and implement fusion kernels using DSL-based approaches (e.g., Triton) enabling fused operations and performance-critical rewrites.
  • Partner with compiler, performance, and accuracy teams to co-design lowering strategies and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
  • Own transformer-specific optimizations including KVcache management, decoding behavior, and long context performance.
  • Enable and optimize continuous batching and dynamic scheduling, impact on memory and tail latency.
  • Architect and scale distributed inference strategies across multi-core and multi-device systems.

Skills

PyTorch
Python
Model optimization
Inference optimization
Graph capture
Transformer architectures
KVcache
Distributed systems

Education

Bachelor's degree in CS/Engineering or related field
MS in Computer Science, ML, CE or EE
PhD in a relevant field (bonus)

Tools

TorchDynamo
Torch.compile
ONNX
Triton

Job description

Qualcomm Technologies, Inc. in Austin, TX, is seeking a Staff Engineer – AI Model Optimization Architect to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm accelerators.

You will collaborate with compiler, performance, and accuracy teams to translate models into accelerator-efficient execution, balancing throughput, latency, memory, and quality across Day0 to production deployment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Model Optimization Architect for LLMs & Multimodal
Staff AI Model Optimization Architect for LLMs & Multimodal

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
AI Hardware Performance Modeling Engineer for LLMs
AI Hardware Performance Modeling Engineer for LLMs

Qualcomm • San Diego (CA)

On-site
USD 162,000 - 243,000
AI Model Optimization Architect
AI Model Optimization Architect

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
Performance Modeling Engineer for AI Accelerators and SoC
Performance Modeling Engineer for AI Accelerators and SoC

Qualcomm Technologies • Austin (TX)

On-site
USD 192,000 - 288,000
AI Hardware Performance Modeling for LLMs
AI Hardware Performance Modeling for LLMs

QUALCOMM, Inc. • San Diego (CA)

On-site
USD 162,000 - 243,000
Annual bonus
RSU grants
Comprehensive benefits
AI Model Optimization Architect
AI Model Optimization Architect

QUALCOMM, Inc. • Austin (TX)

On-site
USD 162,000 - 243,000
ML Engineer: Model Optimization & Edge Deployment
ML Engineer: Model Optimization & Edge Deployment

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Performance Modeling Architect: AI Accelerators & SoCs
Performance Modeling Architect: AI Accelerators & SoCs

QUALCOMM, Inc. • Austin (TX)

On-site
USD 192,000 - 288,000
AI TLM Performance Modeling
AI TLM Performance Modeling

Qualcomm • San Diego (CA)

On-site
USD 162,000 - 243,000
AI TLM Performance Modeling
AI TLM Performance Modeling

QUALCOMM, Inc. • San Diego (CA)

On-site
USD 162,000 - 243,000
Annual bonus
RSU grants
Comprehensive benefits