Senior Principal AI Engineer

Cerence AI

United States

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerence Inc. is seeking an experienced ML Systems Engineer to design and run distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution to maximize throughput and tenacity, while diagnosing bottlenecks across compute, memory, and network. Responsibilities include productionising large‑model training pipelines with PyTorch Distributed, Megatron‑LM and DeepSpeed, and improving training stability at scale.

Qualifications

  • Deep hands-on experience with distributed systems or ML systems.
  • Experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Strong understanding of data, tensor, and pipeline parallelism.
  • Low-level understanding of GPU communication and networking.

Responsibilities

  • Design and operate distributed training systems for large neural networks across GPU clusters.
  • Optimise multi-node, multi-GPU execution to maximise throughput and utilisation.
  • Diagnose and resolve bottlenecks across compute, memory, and network.
  • Improve training stability and fault tolerance at scale.
  • Partner with research and applied ML teams to productionise training pipelines.

Skills

Distributed systems experience
Large‑scale GPU workloads
PyTorch distributed training
Parallelism knowledge (data/tensor/pip
GPU communication & networking
GPU orchestration (Slurm/Kubernetes/R​
NCCL/RDMA/InfiniBand/NVLink
Megatron‑LM/DeepSpeed
Activation checkpointing/ZeRO offload

Tools

Slurm
Kubernetes
Ray
RunAI
NCCL
RDMA
InfiniBand
NVLink

Job description

  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters
  • Optimise multi‑node, multi‑GPU execution to maximize throughput and utilisation
  • Diagnose & resolve bottlenecks across compute, memory, and network
  • Improve training stability and fault tolerance at scale
  • Partner with research and applied ML teams to productionise large‑model training pipelines
A Moving Experience.
What You Will Work On
  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters
  • Optimise multi‑node, multi‑GPU execution to maximise throughput and utilisation
  • Diagnose & resolve bottlenecks across compute, memory, and network
  • Improve training stability and fault tolerance at scale
  • Partner with research and applied ML teams to productionise large‑model training pipelines
Core Responsibilities
  • Distributed Training Infrastructure
  • Build and optimise GPU cluster orchestration using:
  • Slurm
  • Kubernetes
  • Ray
  • RunAI
  • Ensure efficient scheduling, isolation, and fairness across training workloads
  • Communication & Networking
  • Optimize and debug distributed communication using:
  • NCCL
  • RDMA
  • InfiniBand
  • NVLink
  • Minimise networking bottlenecks that dominate end‑to‑end training time
  • Training Frameworks
  • Scale large‑model training using:
  • PyTorch Distributed
  • Megatron‑LM
  • DeepSpeed
  • Own multi‑node launch configurations, failure recovery, and performance tuning
  • Memory & Performance Optimization
  • Apply advanced memory optimisation techniques:
  • Activation checkpointing
  • ZeRO (Stage 1–3) and offload strategies
  • Balance compute, memory, and communication to push model size and batch scale
What Success Looks Like
  • GPU utilisation consistently stays high (>80–90%)
  • Training scales cleanly from single node to dozens or hundreds of GPUs
  • Communication overhead is minimised and predictable
  • Large training jobs run stably for days or weeks without failure
  • New models can be trained faster, larger, and more reliably than before
Required Experience & Skills
  • Strongly Required
  • Deep hands‑on experience with distributed systems or ML systems
  • Experience running large‑scale workloads on GPU clusters
  • Production experience with PyTorch distributed training
  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
  • Low‑level understanding of GPU communication and networking
  • Critical Technical Skills
  • GPU orchestration: Slurm, Kubernetes, Ray, RunAI
  • Communication libraries: NCCL, RDMA, InfiniBand, NVLink
  • Training frameworks: PyTorch Distributed, Megatron‑LM, DeepSpeed
  • Memory optimisation: activation checkpointing, ZeRO offload techniques
Common Problems You’ll Be Solving
  • Many teams fail at scale because:
  • GPU utilisation is low despite large clusters
  • Networking and communication dominate training time
  • Training jobs crash or become unstable at large scale
  • You will be explicitly focused on eliminating these failure modes.
Ideal Background
  • This role is a strong fit for individuals who have worked as:
  • ML Systems Engineer
  • Distributed Systems Engineer
  • AI Infrastructure Engineer
  • HPC Engineer transitioning into ML
  • Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.
Why This Role Matters

Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:

  • Larger models
  • Faster iteration cycles
  • More reliable research‑to‑production pipelines

You will be building the foundation that makes large‑scale AI possible.

Cerence Inc. (Nasdaq: CRNC and www.cerence.com) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.

EQUAL OPPORTUNITY EMPLOYER

Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, colour, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.

  • Following workplace security protocols and training programmes to familiarise with the ways to maintain a safe workplace.
  • Following security procedures to report any suspicious activity.
  • Having respect for corporate security procedures to allow those procedures to be effective.
  • Adhering to company's compliance and regulations.
  • Encouraging to follow a zero tolerance for workplace violence.
  • Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).
  • Demonstrative knowledge of information security through internal training programmes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Principal Software Scientist
Sr. Principal Software Scientist

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus
Insurance coverage
Paid time off
+4
Sr. Principal Software Engineer
Sr. Principal Software Engineer

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Principal Software Engineer – Robot Applications & Voice AI
Principal Software Engineer – Robot Applications & Voice AI

Cerence AI • San Francisco (CA)

Hybrid
USD 166,000 - 264,000
Annual bonus
Medical, dental, vision insurance
Paid time off
+3
Sr. Principal Software Engineer
Sr. Principal Software Engineer

Cerence Inc. • Burlington (MA)

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off
+1
Senior Technical Solution Architect – Industrial AI
Senior Technical Solution Architect – Industrial AI

Cerence AI • United States

Hybrid
USD 132,000 - 212,000
Remote or hybrid work
Annual bonus
Insurance coverage
+3
FinOps Engineer / Architect
FinOps Engineer / Architect

Cerence AI • United States

Hybrid
USD 86,000 - 130,000
Annual bonus opportunity
Insurance coverage (medical, dental,  
Paid time off
+3
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Job stability with startup vitality
Opportunity to work on advanced AI platforms
Open source AI research
Senior Application Developer – Industrial Voice AI & Edge Applications
Senior Application Developer – Industrial Voice AI & Edge Applications

Cerence AI • United States

Hybrid
USD 107,000 - 173,000
Hybrid work model
Competitive compensation
Insurance coverage (medical, dental, &
+1
Business Development Partner, Strategic Accounts & Channel Partners
Business Development Partner, Strategic Accounts & Channel Partners

Cerence Inc. • Northern (KY)

Hybrid
USD 96,000 - 144,000
Insurance coverage (medical, dental, $
Paid time off
Paid holidays
+3
Software Engineer, GPU Inference
Software Engineer, GPU Inference

Cerebras • United States

On-site
USD 150,000 - 210,000