Elite Senior Principal ML Inference Engineer - Edge & CUDA

Cerence AI

United States

Hybrid

USD 185,000 - 280,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Annual bonus
Insurance coverage
Paid time off
401K
Equity awards
Remote/hybrid work

Job summary

Cerence AI is seeking a Senior Principal Software Engineer to advance the future of mobility by optimizing ML inference performance across data center, edge, and embedded platforms.

You will own inference runtimes such as vLLM, TensorRT-LLM, llama.cpp, and QAIRT, while pushing quantization and memory optimizations. This role offers remote or hybrid work options and a competitive equity package.

Qualifications

  • Proven experience optimizing ML inference performance in production.
  • Deep understanding of GPU architecture and memory hierarchies.
  • Hands-on experience with CUDA and low-level performance tuning.
  • Experience deploying models beyond research environments.

Responsibilities

  • Own inference runtimes across data center, edge, and embedded platforms.
  • Optimize and deploy high-performance LLM inference pipelines.
  • Extend and tune inference engines using custom CUDA kernels.
  • Adapt runtimes for constrained and embedded deployment environments.

Skills

ML inference
GPU architecture
CUDA
low-level perf

Tools

vLLM
TensorRT-LLM
llama.cpp
QAIRT

Job description

Cerence AI is seeking a Senior Principal Software Engineer to advance the future of mobility by optimizing ML inference performance across data center, edge, and embedded platforms.

You will own inference runtimes such as vLLM, TensorRT-LLM, llama.cpp, and QAIRT, while pushing quantization and memory optimizations. This role offers remote or hybrid work options and a competitive equity package.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal AI Inference Engineer - Remote/Edge
Senior Principal AI Inference Engineer - Remote/Edge

Cerence Inc. • Burlington (MA)

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off
+1
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Principal ML Engineer — Inference Frameworks (Onsite)
Senior Principal ML Engineer — Inference Frameworks (Onsite)

MemWize, Inc • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
Lead AI Inference & Optimization Engineer (Vehicle & Edge)
Lead AI Inference & Optimization Engineer (Vehicle & Edge)

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior AI Systems Engineer - LLMs, Edge & On-Prem
Senior AI Systems Engineer - LLMs, Edge & On-Prem

Lattice Semiconductor Corp • Nellis Air Force Base Census-Designated Place (NV)

On-site
USD 199,000 - 243,000
Equity compensation
Healthcare and retirement plans
Paid time off
Principal AI Systems Engineer - LLM Inference (Cloud)
Principal AI Systems Engineer - LLM Inference (Cloud)

Qualcomm • San Diego (CA)

On-site
USD 224,000 - 335,000
Senior GPU ML Inference Engineer — Edge AI Platforms
Senior GPU ML Inference Engineer — Edge AI Platforms

NVIDIA • Westford (MA)

On-site
USD 224,000 - 431,250
Senior ML Engineer — Edge AI & Embedded Systems
Senior ML Engineer — Edge AI & Embedded Systems

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits