Distributed AI Training Architect

cerence

United States

On-site

USD 180,000 - 250,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

cerence is seeking an experienced ML infrastructure engineer to design and operate distributed training systems for large neural networks across GPU clusters. You will optimize multi-node, multi-GPU execution and diagnose bottlenecks across compute, memory, and networking to improve training efficiency.

Ideal candidates have hands-on experience with distributed systems, PyTorch distributed training, and deep knowledge of GPU communication.

Qualifications

  • Hands-on experience with distributed or ML systems at scale.
  • Experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Strong understanding of data, tensor, and pipeline parallelism.
  • Low-level understanding of GPU communication and networking.

Responsibilities

  • Design and operate distributed training systems for large neural networks across GPU clusters.
  • Optimize multi-node, multi-GPU execution for throughput and utilization.
  • Diagnose bottlenecks across compute, memory, and network.
  • Improve training stability and fault tolerance at scale.

Skills

Distributed systems
GPU clusters
PyTorch distributed
Parallelism strategies
GPU networking

Tools

Slurm
Kubernetes
Ray
RunAI
NCCL
RDMA
InfiniBand
NVLink
Megatron-LM
DeepSpeed

Job description

cerence is seeking an experienced ML infrastructure engineer to design and operate distributed training systems for large neural networks across GPU clusters. You will optimize multi-node, multi-GPU execution and diagnose bottlenecks across compute, memory, and networking to improve training efficiency.

Ideal candidates have hands-on experience with distributed systems, PyTorch distributed training, and deep knowledge of GPU communication.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Distributed Systems Engineer
Senior AI Distributed Systems Engineer

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000
Senior Principal AI Engineer
Senior Principal AI Engineer

cerence • United States

On-site
USD 180,000 - 250,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Remote Deep Learning Engineer - Distributed Training Scale
Remote Deep Learning Engineer - Distributed Training Scale

BairesDev • Peru (IL)

On-site
USD 120,000 - 180,000
Remote work
Competitive USD compensation
Home setup provided
+3
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500