Senior Distributed AI Training Architect

Luma AI

United States

Remote

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI is seeking an engineer to design and optimize distributed training systems for large multimodal models. You will work on multi-GPU clusters, implement advanced parallelization, and build tools to monitor and debug massive training runs.

Requires extensive distributed PyTorch experience, deep GPU cluster knowledge, and familiarity with NCCL and MPI. Experience with containerization and orchestration is beneficial. This role is a high-depth engineering position at Luma AI.

Qualifications

  • Extensive distributed PyTorch training across foundation-model scales.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with NCCL and MPI for efficient communication.

Responsibilities

  • Design, implement, and optimize distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.

Skills

Distributed PyTorch
FSDP
Tensor Parallel
Pipeline Parallel
Expert Parallel
GPU clusters
Training stability
Resource tuning

Tools

NCCL
MPI
CUDA
Docker
Kubernetes

Job description

Luma AI is seeking an engineer to design and optimize distributed training systems for large multimodal models. You will work on multi-GPU clusters, implement advanced parallelization, and build tools to monitor and debug massive training runs.

Requires extensive distributed PyTorch experience, deep GPU cluster knowledge, and familiarity with NCCL and MPI. Experience with containerization and orchestration is beneficial. This role is a high-depth engineering position at Luma AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000