Distributed AI Training Engineer (GPU Clusters)

lumalabs-ai

San Francisco (CA)

On-site

USD 188,000 - 395,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Luma AI is seeking engineers for the Training Infrastructure team to design and maintain distributed systems that enable training of large-scale multimodal models on thousands of GPUs. You will work with researchers to push the boundaries of AI model development.

Ideal candidates have extensive PyTorch distributed training experience and a deep understanding of GPU clusters, networking, and storage, plus familiarity with NCCL/MPI and Linux scripting.

Qualifications

  • Extensive experience with distributed PyTorch training and model-scale parallelism.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed system optimization.
  • Preferred: strong Linux systems administration and scripting capabilities.
  • Preferred: experience managing training runs across >100 GPUs.
  • Preferred: experience with containerization, orchestration, and cloud infrastructure.

Responsibilities

  • Design, implement, and optimize efficient distributed training systems for models with thousands of GPUs
  • Research and implement advanced parallelization techniques (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel)
  • Build monitoring, visualization, and debugging tools for large-scale training runs
  • Optimize training stability, convergence, and resource utilization across massive clusters

Skills

Distributed PyTorch
GPUs & clusters
NCCL/MPI
Linux scripting
Containerization
Cloud infrastructure

Job description

Luma AI is seeking engineers for the Training Infrastructure team to design and maintain distributed systems that enable training of large-scale multimodal models on thousands of GPUs. You will work with researchers to push the boundaries of AI model development.

Ideal candidates have extensive PyTorch distributed training experience and a deep understanding of GPU clusters, networking, and storage, plus familiarity with NCCL/MPI and Linux scripting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000