Distributed Training Systems Engineer

Reflection AI Ltd

New York (NY)

On-site

USD 180,000 - 270,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited PTO
Sponsorship visas
Team events

Job summary

Reflection AI Ltd in New York seeks a senior distributed training engineer to build and scale systems powering frontier ML model pre-training.

You will design and operate large-scale training runs across thousands of GPUs, developing infrastructure with Megatron, DeepSpeed, and NCCL, and optimizing throughput and memory usage.

Join a mission to make intelligence open and accessible, with stock options, competitive compensation, and visa sponsorship where applicable.

Qualifications

  • Experience building or operating distributed training systems for large ML models.
  • Experience with Megatron, DeepSpeed or similar large-scale training systems.
  • Familiarity with large-scale model parallelism strategies (data, tensor, pipeline, or expert parallelism).
  • Experience optimizing training throughput and GPU utilization in large distributed environments.
  • Familiarity with GPU communication libraries such as NCCL and performance tuning for distributed workloads.
  • Experience working closely with ML researchers to productionize experimental training workflows.
  • Strong debugging skills across GPU compute, distributed training systems, and large-scale ML pipelines.
  • Experience working with large datasets and training pipelines used for foundation model pre-training.

Responsibilities

  • Build and scale distributed training systems for frontier models.
  • Design and operate large-scale training runs across thousands of GPUs.
  • Develop infrastructure for efficient training across thousands of GPUs using Megatron, DeepSpeed, NCCL.
  • Optimize training throughput, stability, and performance of large workloads.
  • Collaborate with researchers to productionize experimental training workflows.
  • Improve GPU utilization and memory management in distributed stacks.
  • Maintain training pipelines for large datasets, checkpoints, and iterations.
  • Debug performance bottlenecks across distributed training stacks, including model parallelism and GPU communication.
  • Contribute to rapid experimentation and iteration on new training techniques.

Skills

Distributed training
Megatron/DeepSpeed
Model parallelism
GPU utilization
Debugging pipelines
Collaboration with researchers

Tools

Megatron
DeepSpeed
NCCL

Job description

Reflection AI Ltd in New York seeks a senior distributed training engineer to build and scale systems powering frontier ML model pre-training.

You will design and operate large-scale training runs across thousands of GPUs, developing infrastructure with Megatron, DeepSpeed, and NCCL, and optimizing throughput and memory usage.

Join a mission to make intelligence open and accessible, with stock options, competitive compensation, and visa sponsorship where applicable.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Systems Engineer - Pre-Training Infra
Staff AI Systems Engineer - Pre-Training Infra

Reflection • New York (NY)

On-site
USD 130,000 - 180,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Senior AI/ML Systems Engineer - Distributed Training
Senior AI/ML Systems Engineer - Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+2
Staff Research Software Engineer: Open ML Infra & RL Training
Staff Research Software Engineer: Open ML Infra & RL Training

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 260,000
Top-tier compensation & equity
Stock options
Health & wellness benefits
+5
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Frontier AI Systems Engineer - Distributed Training (Remote)
Frontier AI Systems Engineer - Distributed Training (Remote)

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Distributed ML Systems Engineer: Scalable Training
Distributed ML Systems Engineer: Scalable Training

Susquehanna International Group, LLP • Lower Merion Township

On-site
USD 120,000 - 170,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500