RL Post-Training Systems Engineer (Distributed Infra)

NVIDIA

New York (NY)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is building an RL post-training infrastructure team to scale experimental workflows from a single GPU to thousands of nodes, delivering reliable, high-performance runtimes for researchers. You will collaborate with researchers and labs, optimize PyTorch-based RL loops, and improve fault tolerance, elastic scaling, and portability across CPU, GPU, and LPUs.

This role offers the chance to contribute to VeRL, Miles, TorchTitan and related ecosystems while shaping production-grade distributed

Qualifications

  • MS or PhD in CS/CE or related field.
  • 5+ years of distributed systems or ML infrastructure experience.
  • Strong Python and C/C++ proficiency.
  • Experience building/operating large-scale distributed systems in production.
  • Excellent verbal and written communication across teams.

Responsibilities

  • Architect and build RL post-training infrastructure at scale from a single GPU to thousands of nodes.
  • Tune RL training–inference–rollout loops for performance on GPUs, CPUs and LPUs.
  • Improve open-source RL frameworks and collaborate with hardware and software teams.
  • Focus on fault tolerance, elastic scaling, and fast restarts for long-running jobs.

Skills

Python
C/C++
Distributed systems
Communication skills

Education

MS or PhD in CS/CE or related field

Tools

PyTorch
Ray
Kubernetes

Job description

NVIDIA is building an RL post-training infrastructure team to scale experimental workflows from a single GPU to thousands of nodes, delivering reliable, high-performance runtimes for researchers. You will collaborate with researchers and labs, optimize PyTorch-based RL loops, and improve fault tolerance, elastic scaling, and portability across CPU, GPU, and LPUs.

This role offers the chance to contribute to VeRL, Miles, TorchTitan and related ecosystems while shaping production-grade distributed

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
RL Post-Training Infrastructure Architect
RL Post-Training Infrastructure Architect

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior Engineering Manager, RL Post-Training Frameworks
Senior Engineering Manager, RL Post-Training Frameworks

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000