Research Software Engineer: Scalable RL Training Systems

Jobzhr

New York (NY)

On-site

USD 180,000 - 240,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Salary and equity
Stock options
Healthcare
Meals provided
Parental leave
Unlimited PTO

Job summary

Reflection is building open AI infrastructure to turn cutting-edge research into scalable training systems. You will design and optimize core training pipelines, RL loops, and large-scale data processing across thousands of GPUs, working closely with researchers to deliver reliable, production-grade infrastructure.

The role emphasizes performance, numerical stability, and reproducibility, with opportunities to contribute across HPC, distributed systems, and advanced frameworks in a high-autonomy

Qualifications

  • Strong software engineer who speaks the language of machine learning.
  • PhD is not required, but you know how to implement a research paper.
  • You have deep experience in at least one: Distributed Training & Inference or Data Infrastructure.
  • You enjoy the boundary between ML, distributed systems and HPC.
  • You care deeply about performance, numerical stability, and reproducibility.

Responsibilities

  • Designing and optimizing large-scale training loops and data pipelines.
  • Implementing state-of-the-art techniques with numerical stability and efficiency.
  • Building internal tools for launching, monitoring, and reproducing experiments.
  • Diagnosing bottlenecks across the training stack (GPU memory, communication, dataloader stalls).
  • Translating research prototypes into production-grade infrastructure.

Skills

Distributed Training
Data Infrastructure
HPC
Numerical Stability
GPU Optimizations

Tools

PyTorch
JAX
Megatron
Triton
Ray

Job description

Reflection is building open AI infrastructure to turn cutting-edge research into scalable training systems. You will design and optimize core training pipelines, RL loops, and large-scale data processing across thousands of GPUs, working closely with researchers to deliver reliable, production-grade infrastructure.

The role emphasizes performance, numerical stability, and reproducibility, with opportunities to contribute across HPC, distributed systems, and advanced frameworks in a high-autonomy

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - Large-Scale GPU Inference & RL Infra
Staff Engineer - Large-Scale GPU Inference & RL Infra

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness benefits
+5
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Research Engineer, Open-Model AI & Scalable Training
Staff Research Engineer, Open-Model AI & Scalable Training

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
Top-tier salary
Stock options
Health insurance
+4
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • New York (NY)

On-site
USD 120,000 - 180,000