AI Model Optimization Architect

Qualcomm

Austin (TX)

On-site

USD 158,000 - 238,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Qualcomm Technologies, Inc. seeks a Staff Engineer to lead end-to-end AI model optimization for LLMs, VLMs, and diffusion on Qualcomm accelerators. You will transform PyTorch models, manage KVcache behavior, and drive deployment with PyTorch, ONNX, and torch.compile across multi-core systems.

You will collaborate with compiler and performance teams to craft lowering strategies and scalable tooling, ensuring high throughput with low latency and production-grade reliability.

Qualifications

  • Bachelor's degree in CS/Engineering with 4+ years of related work experience.
  • Master's degree with 3+ years of related work experience.
  • PhD in CS/Engineering or related field with 2+ years of related work experience.
  • Experience in hardware/software systems and ML accelerators preferred.

Responsibilities

  • Architect and deliver model optimization strategies for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile.
  • Design fusion kernels via DSLs (e.g., Triton) and perform kernel fusion.
  • Collaborate with compiler, performance, and accuracy teams on lowering strategies.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency.
  • Own transformer optimizations including KVcache management and decoding behavior.
  • Support distributed inference strategies across multi-core/multi-device systems.
  • Create reusable patterns and tooling to scale model optimizations to new hardware.
  • Debug complex performance issues to deliver production-ready solutions.

Skills

PyTorch
Python
Transformer architectures
KVcache management
Distributed systems
Model optimization
Research leadership

Education

Bachelor's degree in CS/Engineering
Master's degree in CS/Engineering
PhD in CS/Engineering

Tools

TorchDynamo
Torch.compile
ONNX
Triton

Job description

Company

Qualcomm Technologies, Inc.

Job Area

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary

Qualcomm is leveraging its strengths in compute, connectivity, and AI acceleration to play a central role in the evolution of Cloud AI. The Qualcomm Cloud AI team develops hardware and software platforms enabling efficient inference of large-scale foundation models.

We are seeking a Staff Engineer - AI Model Optimization Architect to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators. This role works closely with compiler, performance, and accuracy teams to translate models into accelerator efficient execution while balancing throughput, latency, memory, and quality. The scope spans Day0 enablement through production deployment, with a strong emphasis on scaling optimizations to future architectures.

Key Responsibilities
  • Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
  • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
  • Partner deeply with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
  • Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
  • Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency.
  • Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
  • Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling.
  • Debug complex performance or stability issues to root cause and drive production ready solutions.
Required Qualifications
  • Expert level expertise in PyTorch and inference focused model optimization; strong Python engineering skills.
  • Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
  • Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
  • Strong foundation in computer architecture, ML accelerators, and distributed systems.
  • Proven ability to lead cross-functional technical efforts and influence design decisions.
  • MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.
Preferred / Bonus Qualifications
  • Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
  • Familiarity with LLM serving stacks and continuous batching systems.
  • Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
  • PhD in a relevant field.
Minimum Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • OR Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • OR PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.

Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail disability-accomodations@qualcomm.com or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities.

EEO Employer: Qualcomm is an equal opportunity employer; all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or any other protected classification.

Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

Pay range and Other Compensation & Benefits

$158,400.00 - $237,600.00

The above pay scale reflects the broad, minimum to maximum, pay scale for this job code for the location for which it has been posted. Even more importantly, please note that salary is only one component of total compensation at Qualcomm. We also offer a competitive annual discretionary bonus program and opportunity for annual RSU grants (employees on sales-incentive plans are not eligible for our annual bonus).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Model Optimization Architect
AI Model Optimization Architect

QUALCOMM, Inc. • Austin (TX)

On-site
USD 162,000 - 243,000
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff

Qualcomm • San Diego (CA)

On-site
USD 178,400 - 267,600
Competitive annual discretionary bonus
RSU grants
Highly competitive benefits package
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)

Qualcomm • Raleigh (NC)

On-site
USD 162,000 - 243,000
Machine Learning Engineer, Staff (Model Optimization)
Machine Learning Engineer, Staff (Model Optimization)

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
AI TLM Performance Modeling
AI TLM Performance Modeling

QUALCOMM, Inc. • San Diego (CA)

On-site
USD 162,000 - 243,000
Annual bonus
RSU grants
Comprehensive benefits
Machine Learning Engineer, Staff (Model Optimization) San Diego, California, United States of America Machine Learning Engineering
Machine Learning Engineer, Staff (Model Optimization) San Diego, California, United States of America Machine Learning Engineering

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Staff/Sr. Staff Software Engineer, AI Software Tools Development
Staff/Sr. Staff Software Engineer, AI Software Tools Development

Qualcomm • San Diego (CA)

On-site
USD 158,400 - 237,600
Annual discretionary bonus
RSU grants
Comprehensive benefits package
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)
Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Senior Machine Learning Engineer – AI/ML Compiler
Senior Machine Learning Engineer – AI/ML Compiler

Qualcomm • Santa Clara (CA)

On-site
USD 151,000 - 227,000
Staff Machine Learning Engineer – AI/ML Compiler
Staff Machine Learning Engineer – AI/ML Compiler

Qualcomm • Santa Clara (CA)

On-site
USD 160,500 - 240,700