Senior AI Distributed Systems Engineer

United States Digital Space LLC

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerence Inc. is seeking an ML Infrastructure Engineer to design and operate distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution and diagnose bottlenecks in compute, memory, and networking to maximize throughput and training stability. You will partner with research and applied ML teams to productionize large-model training pipelines using PyTorch Distributed, Megatron-LM, and DeepSpeed, with emphasis on scalable,

Qualifications

  • Strong hands-on experience with distributed ML systems.
  • Experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Understanding of parallelism strategies (data, tensor, pipeline).

Responsibilities

  • Design and operate distributed training systems for large neural networks across GPU clusters.
  • Optimize multi-node, multi-GPU execution for throughput and utilization.
  • Diagnose bottlenecks across compute, memory, and network.
  • Improve training stability and fault tolerance at scale.
  • Collaborate with R&D teams to productionize training pipelines.

Skills

Distributed systems
Large-model training
GPU clusters
PyTorch Distributed
Megatron-LM
DeepSpeed
NCCL
Kubernetes
Slurm
Ray

Tools

Slurm
Kubernetes
Ray
RunAI
NCCL
RDMA
InfiniBand
NVLink
PyTorch Distributed
Megatron-LM
DeepSpeed

Job description

Cerence Inc. is seeking an ML Infrastructure Engineer to design and operate distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution and diagnose bottlenecks in compute, memory, and networking to maximize throughput and training stability. You will partner with research and applied ML teams to productionize large-model training pipelines using PyTorch Distributed, Megatron-LM, and DeepSpeed, with emphasis on scalable,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000
Distributed AI Training Architect
Distributed AI Training Architect

cerence • United States

On-site
USD 180,000 - 250,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Senior Principal AI Engineer
Senior Principal AI Engineer

cerence • United States

On-site
USD 180,000 - 250,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Senior Principal AI Engineer
Senior Principal AI Engineer

Cerence AI • United States

On-site
USD 180,000 - 240,000