Research Scientist / Engineer – Training Infrastructure

Luma

Redwood City (CA)

On-site

USD 210,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma is seeking an engineer to design and optimize distributed training systems for multimodal models across thousands of GPUs. You will tackle advanced parallelism, training stability, and utilization at scale.

You will work with PyTorch, CUDA, NCCL, and Tensor Parallel, building monitoring and tooling for large-scale runs. This role suits someone who has shipped foundation-model training at scale and can improve stability and efficiency on massive clusters.

Qualifications

  • Experience in extensive distributed PyTorch training and parallelisms for foundation-models.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization.

Responsibilities

  • Design, implement, and optimize distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.
  • Immerse & diagnose training stacks, land a parallelization or stability improvement, and scale tooling for reliability.

Skills

Distributed PyTorch training
GPU clusters & networking
Distributed-system optimization

Tools

NCCL
MPI

Job description

You’ll build the distributed systems that train Luma's large-scale multimodal models across thousands of GPUs, so researchers can focus on innovation on top of reliable, efficient, scalable infrastructure.

This is hard PyTorch, CUDA, and distributed-systems work - advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models at scale. If you haven't worked at the level of FSDP and multi-node training, this is the wrong depth.

What You’ll Own
  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.
First 90 Days
  • Days 1-30 - Immerse & Diagnose: Learn the current training stack and where stability and utilization hurt at scale.
  • Days 30-60 - Ship & Validate: Land a parallelization or stability improvement that measurably helps a real training run.
  • Days 60-90 - Scale & Systemize: Build the monitoring and tooling that keeps large runs reliable and efficient.
What You Bring
  • Extensive distributed PyTorch training and parallelisms in foundation-model training.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization.
Nice to Have
  • Strong Linux systems administration and scripting.
  • Experience managing training runs across 100+ GPUs.
  • Experience with containerization, orchestration, and cloud infrastructure.

About Luma: Luma’s mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence - the next step beyond language models comes from vision. Luma is an equal opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000