Staff Engineer: Distributed ML Training Systems

reflectionai

San Francisco, New York (CA, NY)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options
Health insurance
Meals provided
Parental leave
Unlimited vacation
Visa sponsorship
Team off-sites

Job summary

reflectionai seeks a senior engineer to build and scale distributed training systems powering frontier model pre-training in San Francisco. You will partner with research teams to run large-scale foundation model training across thousands of GPUs, designing infrastructure for efficient, production-ready workflows.

You will optimize throughput, memory, and GPU utilization while debugging complex distributed pipelines and collaborating with researchers to push novel training techniques into

Qualifications

  • Experience building or operating distributed training systems for large ML models.
  • Strong experience with Megatron, DeepSpeed, or similar frameworks.
  • Familiarity with model parallelism strategies (data/tensor/pipeline/expert).
  • Experience optimizing throughput and GPU utilization in large distributed envs.
  • Familiarity with NCCL and performance tuning for distributed workloads.
  • Experience collaborating with ML researchers to productionize training workflows.
  • Strong debugging across GPU compute, distributed systems, and ML pipelines.
  • Experience handling large datasets and training pipelines for foundation models.

Responsibilities

  • Build and scale distributed training systems that power frontier model pre-training.
  • Work with research teams to run large-scale training for foundation models.
  • Develop infra enabling training across thousands of GPUs using modern distributed frameworks.
  • Optimize training throughput, stability, and efficiency for large workloads.
  • Collaborate with researchers to translate ideas into production-ready training systems.
  • Improve performance via optimization of communication, memory, and GPU utilization.
  • Build and maintain training pipelines for large datasets, checkpoints, and iterations.
  • Debug and resolve bottlenecks across distributed training stacks and model parallelism.

Skills

Distributed training
Megatron/DeepSpeed
GPU utilization
NCCL
Research collaboration
Performance profiling
Large-scale pipelines

Tools

Megatron
DeepSpeed
NCCL

Job description

reflectionai seeks a senior engineer to build and scale distributed training systems powering frontier model pre-training in San Francisco. You will partner with research teams to run large-scale foundation model training across thousands of GPUs, designing infrastructure for efficient, production-ready workflows.

You will optimize throughput, memory, and GPU utilization while debugging complex distributed pipelines and collaborating with researchers to push novel training techniques into

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Staff AI Systems Engineer - Pre-Training Infra
Staff AI Systems Engineer - Pre-Training Infra

Reflection • New York (NY)

On-site
USD 130,000 - 180,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: High-Performance Distributed Training
ML Systems Engineer: High-Performance Distributed Training

Motional • San Francisco (CA)

Hybrid
USD 144,000 - 192,000
Medical
Dental
Vision
+4
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff, Pre-Training Infra — Distributed ML Training
Staff, Pre-Training Infra — Distributed ML Training

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+2
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Senior ML Infrastructure Engineer — Frontier RL & LLM Training
Senior ML Infrastructure Engineer — Frontier RL & LLM Training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1