Senior Distributed ML Training Engineer

Visa Hunt

San Francisco, Northern (CA, KY)

Hybrid

USD 220,000 - 310,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Health/dental/vision insurance
Meals provided in office
Parental leave (22+ weeks)

Job summary

Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to design scalable training runs across thousands of GPUs, optimizing throughput and stability while advancing production-ready training pipelines.

This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving communication and memory efficiency in large-scale training environments.

Qualifications

  • Experience building or operating distributed training systems for large ML models.
  • Strong experience with Megatron, DeepSpeed, or similar large-scale training frameworks.
  • Familiarity with large-scale model parallelism strategies (data, tensor, pipeline, or expert).
  • Experience optimizing training throughput and GPU utilization in large environments.
  • Proven debugging skills across GPU compute and distributed training stacks.

Responsibilities

  • Build and scale distributed training systems for frontier model pre-training.
  • Operate large-scale training runs with research teams.
  • Develop infrastructure for training across thousands of GPUs.
  • Optimize throughput, stability, and efficiency of training workloads.
  • Collaborate with researchers to productionize training workflows.
  • Improve GPU communication and memory usage in distributed stacks.
  • Maintain training pipelines with large datasets and checkpoints.
  • Debug and resolve performance bottlenecks in model parallelism and runtime systems.

Skills

Distributed training systems
Megatron
DeepSpeed
Model parallelism
GPU optimization
NCCL
Debugging distributed stacks

Tools

Megatron
DeepSpeed
NCCL

Job description

Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to design scalable training runs across thousands of GPUs, optimizing throughput and stability while advancing production-ready training pipelines.

This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving communication and memory efficiency in large-scale training environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior ML Engineer, Distributed Training & P2P Systems
Senior ML Engineer, Distributed Training & P2P Systems

Pluralis Research • California (MO)

Remote
USD 120,000 - 160,000
Staff Engineer - Large-Scale GPU Inference & RL Infra
Staff Engineer - Large-Scale GPU Inference & RL Infra

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness benefits
+5
Senior Distributed Training Engineer — Remote & Flexible Hours
Senior Distributed Training Engineer — Remote & Flexible Hours

BairesDev • Peru (IL)

On-site
USD 120,000 - 180,000
100% remote work
Competitive USD compensation
Hardware provided
+2
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000