Senior AI Systems Architect: Scalable Distributed Training

Cerence AI

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cerence Inc. is seeking an experienced ML Systems Engineer to design and run distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution to maximize throughput and tenacity, while diagnosing bottlenecks across compute, memory, and network. Responsibilities include productionising large‑model training pipelines with PyTorch Distributed, Megatron‑LM and DeepSpeed, and improving training stability at scale.

Qualifications

  • Deep hands-on experience with distributed systems or ML systems.
  • Experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Strong understanding of data, tensor, and pipeline parallelism.
  • Low-level understanding of GPU communication and networking.

Responsibilities

  • Design and operate distributed training systems for large neural networks across GPU clusters.
  • Optimise multi-node, multi-GPU execution to maximise throughput and utilisation.
  • Diagnose and resolve bottlenecks across compute, memory, and network.
  • Improve training stability and fault tolerance at scale.
  • Partner with research and applied ML teams to productionise training pipelines.

Skills

Distributed systems experience
Large‑scale GPU workloads
PyTorch distributed training
Parallelism knowledge (data/tensor/pip
GPU communication & networking
GPU orchestration (Slurm/Kubernetes/R​
NCCL/RDMA/InfiniBand/NVLink
Megatron‑LM/DeepSpeed
Activation checkpointing/ZeRO offload

Tools

Slurm
Kubernetes
Ray
RunAI
NCCL
RDMA
InfiniBand
NVLink

Job description

Cerence Inc. is seeking an experienced ML Systems Engineer to design and run distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution to maximize throughput and tenacity, while diagnosing bottlenecks across compute, memory, and network. Responsibilities include productionising large‑model training pipelines with PyTorch Distributed, Megatron‑LM and DeepSpeed, and improving training stability at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Principal AI Engineer Distributed Training Architect
Senior Principal AI Engineer Distributed Training Architect

Cerence Inc. • United States

Remote
USD 180,000 - 260,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
AI Training Systems Architect (Distributed)
AI Training Systems Architect (Distributed)

Unconventional AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Health benefits
401k matching
Unlimited PTO
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500