AI Model Optimization Architect

Qualcomm

San Diego (CA)

On-site

USD 158,400 - 237,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive annual discretionary bonus
Annual RSU grants
Highly competitive benefits package

Job summary

Qualcomm is seeking a Staff Engineer – AI Model Optimization Architect in San Diego, California. This position involves leading end-to-end model transformation and optimization for large-scale models on Qualcomm accelerators, collaborating with compiler and performance teams to enhance execution efficiency.

The ideal candidate will have 4+ years of relevant experience, expert-level skills in PyTorch, and a strong background in computer architecture and ML accelerators. The role offers a competitive salary of $158,400 to $237,600 along with a robust benefits package.

Qualifications

  • 4+ years of experience in Hardware Engineering, Software Engineering, or related.
  • Hands-on experience with torch.compile or related workflows.
  • Strong foundation in computer architecture and ML accelerators.

Responsibilities

  • Lead end-to-end model optimization for LLMs and multimodal models.
  • Architect model optimization strategies for Qualcomm accelerators.
  • Debug complex performance issues and drive solutions.

Skills

Expert level expertise in PyTorch
Strong Python engineering skills
Deep understanding of transformer architectures
Knowledge of graph capture and compilation workflows

Education

MS in Computer Science, Machine Learning or related
PhD in relevant field

Tools

Triton

Job description

Company

Qualcomm Technologies, Inc.

Job Area

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary

Qualcomm is leveraging its strengths in compute, connectivity, and AI acceleration to play a central role in the evolution of Cloud AI. The Qualcomm Cloud AI team develops hardware and software platforms enabling efficient inference of large-scale foundation models.

Position Overview

We are seeking a Staff Engineer – AI Model Optimization Architect to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators. This role works closely with compiler, performance, and accuracy teams to translate models into accelerator efficient execution while balancing throughput, latency, memory, and quality. The scope spans Day0 enablement through production deployment, with a strong emphasis on scaling optimizations to future architectures.

Key Responsibilities
  • Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
  • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
  • Partner deeply with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
  • Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
  • Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency.
  • Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
  • Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling.
  • Debug complex performance or stability issues to root cause and drive production ready solutions.
Required Qualifications
  • Expert level expertise in PyTorch and inference focused model optimization; strong Python engineering skills.
  • Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
  • Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
  • Strong foundation in computer architecture, ML accelerators, and distributed systems.
  • Proven ability to lead cross-functional technical efforts and influence design decisions.
  • MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.
Preferred / Bonus Qualifications
  • Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
  • Familiarity with LLM serving stacks and continuous batching systems.
  • Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
  • PhD in a relevant field.
Minimum Qualifications

Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.

Pay Range and Benefits

$158,400.00 – $237,600.00. Salary is one component of total compensation. Competitive annual discretionary bonus program, opportunity for annual RSU grants. Highly competitive benefits package. For more details, refer to Qualcomm U.S. benefits information.

Equal Opportunity Employer

Qualcomm is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or any other protected classification. Qualcomm is committed to providing reasonable accommodations for individuals with disabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff

Qualcomm • San Diego (CA)

On-site
USD 178,400 - 267,600
Competitive annual discretionary bonus
RSU grants
Highly competitive benefits package
Sr Software Engineer, AI Tools – On-Device Generative AI Model Optimization
Sr Software Engineer, AI Tools – On-Device Generative AI Model Optimization

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive benefits package
Annual RSU grants
Discretionary bonus program
Staff Machine Learning Engineer – AI/ML Compiler
Staff Machine Learning Engineer – AI/ML Compiler

Qualcomm • San Diego (CA)

On-site
USD 160,500 - 240,700
Annual discretionary bonus program
Opportunity for RSU grants
Comprehensive benefits package
Sr. AI Engineer
Sr. AI Engineer

Qualcomm • Santa Clara (CA)

On-site
USD 129,000 - 194,000
Comprehensive healthcare
Retirement plans
Annual discretionary bonus program
+1
Senior Engineer - Machine Learning
Senior Engineer - Machine Learning

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive annual discretionary bonus program
Opportunity for annual RSU grants
Comprehensive benefits package
Staff Machine Learning Engineer – AI/ML Compiler
Staff Machine Learning Engineer – AI/ML Compiler

Qualcomm • Santa Clara (CA)

On-site
USD 160,500 - 240,700
Staff Machine Learning Engineer – Model Optimization & Quantization
Staff Machine Learning Engineer – Model Optimization & Quantization

Socket.dev • Santa Clara (CA)

On-site
USD 161,000 - 241,000
Software Engineer, AI Tools – Delegate
Software Engineer, AI Tools – Delegate

Qualcomm • Raleigh (NC)

On-site
USD 110,000 - 166,000
Staff AI Model Optimization Architect
Staff AI Model Optimization Architect

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus
Annual RSU grants
Highly competitive benefits package
Staff/Sr. Staff Software Engineer, AI Software Tools Development
Staff/Sr. Staff Software Engineer, AI Software Tools Development

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Annual discretionary bonus
RSU grants
Comprehensive benefits package