Senior RL Post-Training Systems Engineer

NVIDIA

Massachusetts

On-site

USD 224,000 - 357,000

Full time

29 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is building an RL Frameworks engineering team to develop open-source tools and infrastructure for AI researchers and post-training teams. The role spans the full software stack, from researchers to distributed runtimes, including VeRL, Miles, TorchTitan, Ray, Monarch, and more.

You will architect and build RL post-training infrastructure that scales from single-GPU experiments to production across thousands of nodes, tuning loops for peak performance while improving framework usability

Qualifications

  • MS or PhD in Computer Science, Computer Engineering, or related field (or equivalent).
  • 5+ years in distributed systems, HPC, or ML systems engineering.
  • Strong Python and C/C++ proficiency.
  • Experience building large-scale distributed systems in production.
  • Excellent verbal and written communication across teams.

Responsibilities

  • Architect and build RL post-training infrastructure that scales from single-GPU to thousands of nodes.
  • Tune training-inference-rollout loops on GPUs, CPUs, and LPUs for performance.
  • Improve open-source RL frameworks and partner with hardware and runtime teams.
  • Ensure fault tolerance, elastic scaling, and fast restarts for long-running jobs.
  • Collaborate with researchers and partners to prioritize RL workload capabilities.

Skills

Python
C/C++
Distributed systems
Communication
Research experience

Education

MS or PhD in CS/CE or equivalent

Tools

PyTorch
Kubernetes
Ray
Monarch

Job description

NVIDIA is building an RL Frameworks engineering team to develop open-source tools and infrastructure for AI researchers and post-training teams. The role spans the full software stack, from researchers to distributed runtimes, including VeRL, Miles, TorchTitan, Ray, Monarch, and more.

You will architect and build RL post-training infrastructure that scales from single-GPU experiments to production across thousands of nodes, tuning loops for peak performance while improving framework usability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Post-Training Frameworks Architect
RL Post-Training Frameworks Architect

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Research Software Engineer: Scalable RL Training Systems
Research Software Engineer: Scalable RL Training Systems

Jobzhr • New York (NY)

On-site
USD 180,000 - 240,000
Salary and equity
Stock options
Healthcare
+3
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000