RL Post-Training Infrastructure Architect

NVIDIA

Washington (District of Columbia)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is building an RL Frameworks engineering team to develop the open-source tools and infrastructure that AI researchers and post-training teams depend on. The team spans the full software stack, from collaborating closely with researchers to improving runtimes like Ray and Monarch.

You will architect and build RL post-training infrastructure that scales from a single GPU to production across thousands of nodes, tuning loops for performance and reliability across GPUs, CPUs, and LPUs.

Qualifications

  • MS or PhD in CS/CE or related field or equivalent experience.
  • 5+ years of distributed systems or ML infrastructure experience.
  • Strong Python and C/C++ proficiency.
  • Excellent verbal and written communication across teams.

Responsibilities

  • Architect and build RL post-training infrastructure at scale.
  • Tune RL training-inference-rollout loops across GPUs/CPUs/LPUs.
  • Improve open-source RL frameworks and distributed runtimes.

Skills

Python
C/C++
Distributed systems
Communication skills

Education

MS or PhD in Computer Science / Computer Engineering or related field

Tools

VeRL
Miles
TorchTitan
Ray
Monarch

Job description

NVIDIA is building an RL Frameworks engineering team to develop the open-source tools and infrastructure that AI researchers and post-training teams depend on. The team spans the full software stack, from collaborating closely with researchers to improving runtimes like Ray and Monarch.

You will architect and build RL post-training infrastructure that scales from a single GPU to production across thousands of nodes, tuning loops for performance and reliability across GPUs, CPUs, and LPUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
RL Post-Training Systems Engineer (Distributed Infra)
RL Post-Training Systems Engineer (Distributed Infra)

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior Engineering Manager, RL Post-Training Frameworks
Senior Engineering Manager, RL Post-Training Frameworks

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
Senior RL Post-Training Frameworks Leader
Senior RL Post-Training Frameworks Leader

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Head of RL Post-Training Frameworks & Systems
Head of RL Post-Training Frameworks & Systems

NVIDIA • Santa Clara (CA)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3