Senior Principal AI Inference Engineer (Remote)

Cerence

United States

Hybrid

USD 185,000 - 280,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Annual bonus opportunity
Insurance coverage (medical, dental,…)
Paid time off
Paid holidays
RRSP contribution
Equity awards
Remote/hybrid work options

Job summary

Cerence Inc. is seeking a Senior Principal Software Engineer in the United States to lead optimization of ML inference pipelines across data center, edge, and embedded targets.

You will drive quantisation, cache optimization, and runtime improvements to deliver low latency and high throughput for Cerence’s AI-powered mobility solutions. The role demands deep GPU hardware knowledge, hands-on CUDA development, and proven production deployment experience, with a focus on scalable,

Qualifications

  • Proven experience optimizing ML inference performance in production.
  • Deep understanding of GPU architecture and memory hierarchies.
  • Hands-on experience with CUDA and low-level performance tuning.
  • Experience deploying models beyond research environments.

Responsibilities

  • Optimize and deploy high-performance ML inference pipelines.
  • Own inference runtimes across data center, edge, and embedded platforms.
  • Push model performance through quantisation, kernel fusion, and cache optimization.
  • Drive latency and throughput improvements that directly impact production products.
  • Enable efficient, reliable deployment without external vendor dependency.

Skills

ML inference optimization
GPU architecture
CUDA
Low-level performance tuning
Production deployment

Tools

vLLM
TensorRT-LLM
llama.cpp
QAIRT

Job description

Cerence Inc. is seeking a Senior Principal Software Engineer in the United States to lead optimization of ML inference pipelines across data center, edge, and embedded targets.

You will drive quantisation, cache optimization, and runtime improvements to deliver low latency and high throughput for Cerence’s AI-powered mobility solutions. The role demands deep GPU hardware knowledge, hands-on CUDA development, and proven production deployment experience, with a focus on scalable,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal AI Inference Engineer - Remote/Edge
Senior Principal AI Inference Engineer - Remote/Edge

Cerence Inc. • Burlington (MA)

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off
+1
Senior LLM Inference Architect — Edge, Data Center, Remote
Senior LLM Inference Architect — Edge, Data Center, Remote

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Sr. Principal Software Engineer
Sr. Principal Software Engineer

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
Senior AI Infrastructure Engineer, Inference & Optimization
Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Lead AI Inference & Optimization Engineer (Vehicle & Edge)
Lead AI Inference & Optimization Engineer (Vehicle & Edge)

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior Principal AI Engineer Distributed Training Architect
Senior Principal AI Engineer Distributed Training Architect

Cerence Inc. • United States

Remote
USD 180,000 - 260,000
Senior Principal Engineer - Inference Cloud
Senior Principal Engineer - Inference Cloud

Cerebras Systems • Sunnyvale (CA)

On-site
USD 230,000 - 340,000