Sr Software Engineer, AI Tools – On-Device Generative AI Model Optimization

Qualcomm

San Diego (CA)

On-site

USD 140,800 - 211,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive benefits package
Annual RSU grants
Discretionary bonus program

Job summary

Qualcomm is seeking a Machine Learning Engineer in San Diego, CA, to develop innovative AI tools and optimize models for next-generation technologies. This role focuses on model reauthoring and inference optimization for Qualcomm hardware.

The ideal candidate will have 4+ years of experience in software and ML engineering, with proficiency in Python and strong communication skills. Join Qualcomm to shape the future of AI technology!

Qualifications

  • 4+ years of Software Engineering or ML Engineering experience.
  • Experience optimizing model performance for edge hardware.
  • Proficient in Python within large codebases.

Responsibilities

  • Develop and implement cutting-edge AI solutions.
  • Optimize generative AI architectures for Qualcomm hardware.
  • Collaborate cross-functionally with teams on optimization strategies.

Skills

Software Engineering
Machine Learning
Python
Communications

Education

Bachelor’s in Computer Science or related field
Master’s in Computer Science or related field
PhD in Computer Science or related field

Tools

PyTorch
HuggingFace transformers

Job description

Company

Qualcomm Technologies, Inc.

Job Area

Engineering Group: Machine Learning Engineering

General Summary

As a leading technology innovator, Qualcomm pushes the boundaries of what’s possible to enable next‑generation AI experiences and drive agentic transformation, creating a smarter, connected future for all. As a Qualcomm Machine Learning Engineer, you will develop and implement cutting‑edge tools and solutions to enable state‑of‑the‑art AI solutions across various technology verticals.

All Qualcomm employees are expected to actively support diversity on their teams and within the Company.

Location

This role is open to both San Diego, CA and Raleigh, NC and will be onsite full‑time.

What You’ll Do
Model Reauthoring & Architecture Adaptation
  • Reauthor generative AI architectures for efficient execution on Qualcomm AI hardware. This covers LLMs (Llama, Phi, Qwen) and multimodal models (vision‑language, speech, diffusion), including custom attention, normalization, positional embedding, and modality‑specific components.
  • Translate hardware execution constraints — operator support, memory layout, dispatch behavior — into model‑level transformations. These transformations need to preserve accuracy while enabling efficient on‑device execution.
  • Build clean extension points so internal teams and external contributors can onboard new architectures without changing core pipeline code.
Inference Optimization for Edge Hardware
  • Integrate inference acceleration techniques into the model preparation pipeline. This includes memory‑efficient attention, decode acceleration, and serving‑time optimizations.
  • Translate end‑customer deployment constraints — target SoC, context length, latency budget, memory envelope — into concrete model preparation strategies.
Custom Model & OEM Enablement
  • Work with research teams to develop reauthoring strategies for custom OEM models and customer‑specific use cases. Take research prototypes and turn them into production deployments.
Cross‑Functional Collaboration
  • Partner with compiler teams to understand on‑target constraints. Decide on the right response: a graph‑level optimization or model‑level reauthoring.
  • Partner with quantization engineers so architectural decisions compose cleanly with the quantization stack.
Pipeline & Tooling
  • Contribute reauthoring and adaptation stages to a multi‑stage model preparation pipeline. Build developer‑facing diagnostics that give clear, actionable feedback when models fail to lower or run efficiently.
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or related field and 4+ years of Software Engineering, ML Engineering, or related experience.
  • OR Master’s degree in Computer Science, Engineering, or related field and 3+ years of relevant experience.
  • OR PhD in Computer Science, Engineering, or related field and 2+ years of relevant experience.
  • 2+ years in ML systems, model optimization, or inference engineering.
  • Proficient in Python in large, typed codebases.
  • Strong written and verbal communication. Comfortable operating across compiler, research, and partner‑facing teams.
Preferred Qualifications
  • Deep implementation‑level knowledge of generative AI architectures across LLMs and multimodal models.
  • Demonstrated experience optimizing inference for edge or resource‑constrained deployments, with measurable latency or memory wins to point to.
  • Strong PyTorch internals knowledge — module customization, export flows, tracing. Familiarity with the HuggingFace transformers ecosystem.
  • Familiarity with on‑device runtimes and SoC‑level constraints (memory bandwidth, compute precision, NPU/DSP execution). Exposure to QAIRT/QNN, ONNXRuntime, LiteRT‑LLM or similar is a plus.
  • Working understanding of how quantization interacts with model architecture decisions, even if you’re not a quantization specialist.
  • Experience using agentic coding tools such as GitHub Copilot, Cursor, Claude Code, Codeium, or similar AI‑assisted development tools to improve coding productivity and problem‑solving.
Level of Responsibility
  • Works independently on open‑ended optimization challenges. Provides technical guidance and mentorship to teammates.
  • Decisions have broad impact on model accuracy, on‑device performance, and the developer experience of teams using the preparation pipeline.
  • Communicates complex model architecture and inference optimization concepts to a range of audiences: hardware engineers, research scientists, compiler engineers, OEM partners, and external developers.
  • Has meaningful influence on the generative AI optimization roadmap, supported model strategy, and cross‑team integration priorities.
Pay range and Other Compensation & Benefits

$140,800.00 - $211,200.00

The above pay scale reflects the broad, minimum to maximum, pay scale for this job code for the location for which it has been posted. Even more importantly, please note that salary is only one component of total compensation at Qualcomm. We also offer a competitive annual discretionary bonus program and opportunity for annual RSU grants (employees on sales‑incentive plans are not eligible for our annual bonus). In addition, our highly competitive benefits package is designed to support your success at work, at home, and at play.

Equal Opportunity

Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Engineer, Machine Learning (On-Device Software)
Sr Engineer, Machine Learning (On-Device Software)

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Annual discretionary bonus
RSU grants
Competitive benefits package
Sr Software Engineer, AI Tools – AI/ML Compiler
Sr Software Engineer, AI Tools – AI/ML Compiler

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive annual discretionary bonus
Annual RSU grants
Highly competitive benefits package
Senior AI Research Quantization Engineer
Senior AI Research Quantization Engineer

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive annual discretionary bonus
Opportunity for annual RSU grants
Software Engineer, Core AI Software
Software Engineer, Core AI Software

Qualcomm • San Diego (CA)

On-site
USD 122,000 - 185,000
Senior Engineer - Machine Learning
Senior Engineer - Machine Learning

Qualcomm • San Diego (CA)

On-site
USD 140,000 - 212,000
Competitive annual discretionary bonus program
Opportunity for annual RSU grants
Comprehensive benefits package
Entry level & Senior Software Engineer, Core AI Software (Onsite)
Entry level & Senior Software Engineer, Core AI Software (Onsite)

Qualcomm • San Diego (CA)

On-site
USD 140,800 - 211,200
Comprehensive benefits package
Annual discretionary bonus program
Annual RSU grants
Staff Software Engineer, Core AI Software (Mobile BU Focus) _ Onsite
Staff Software Engineer, Core AI Software (Mobile BU Focus) _ Onsite

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus
Annual RSU grants
Comprehensive benefits package
Staff & Senior Staff Software Engineer, AI Software Platform (Onsite)
Staff & Senior Staff Software Engineer, AI Software Platform (Onsite)

Qualcomm • San Diego (CA)

On-site
USD 158,400 - 237,600
Annual discretionary bonus
Competitive benefits package
Opportunity for RSU grants
Staff & Senior Staff Software Engineer, Core AI Software (Onsite)
Staff & Senior Staff Software Engineer, Core AI Software (Onsite)

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus
Annual RSU grants
Comprehensive benefits package
Sr Engineer, Machine Learning Engineering (ML Apps)
Sr Engineer, Machine Learning Engineering (ML Apps)

Qualcomm • San Diego (CA)

On-site
USD 140,800 - 211,200
Competitive annual discretionary bonus
Annual RSU grants
Comprehensive benefits package