Senior RL Post-Training Systems Engineer

NVIDIA

Indiana (PA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a skilled RL frameworks engineer to architect and build scalable post-training infrastructure. You will optimize the RL training‑inference‑rollout loop across GPUs, CPUs, and LPUs, and contribute to open‑source RL frameworks and distributed runtimes.

You will collaborate with researchers, software and hardware teams to ensure the platforms meet the needs of cutting‑edge AI research, deliver reliable performance, and support production‑scale workloads across diverse environments.

Qualifications

  • MS or PhD in Computer Science, Computer Engineering, or related field (or equivalent experience).
  • 5+ years of professional experience in distributed systems, HPC, or ML infrastructure.
  • Strong proficiency in Python and C/C++.

Responsibilities

  • Architect and build RL post-training infrastructure that scales from experimentation on a single GPU to production across thousands of nodes.
  • Tune RL training‑inference‑rollout loops on GPUs, CPUs, and LPUs for performance where it matters.
  • Improve open-source RL frameworks and distributed runtimes (e.g., VeRL, Miles, TorchTitan, Ray, Monarch).
  • Collaborate with researchers and hardware teams to prioritize capabilities and deliver robust production platforms.

Skills

Python
C/C++
Distributed systems
High-performance computing
Communication

Education

MS or PhD in Computer Science/Engineering
Equivalent experience

Tools

PyTorch
Kubernetes
Ray
Monarch
TensorRT

Job description

NVIDIA is seeking a skilled RL frameworks engineer to architect and build scalable post-training infrastructure. You will optimize the RL training‑inference‑rollout loop across GPUs, CPUs, and LPUs, and contribute to open‑source RL frameworks and distributed runtimes.

You will collaborate with researchers, software and hardware teams to ensure the platforms meet the needs of cutting‑edge AI research, deliver reliable performance, and support production‑scale workloads across diverse environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Lead RL Infrastructure Engineer — Scalable GPU Training
Lead RL Infrastructure Engineer — Scalable GPU Training

AMD • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Competitive benefits package
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Research Software Engineer: Scalable RL Training Systems
Research Software Engineer: Scalable RL Training Systems

Jobzhr • New York (NY)

On-site
USD 180,000 - 240,000
Salary and equity
Stock options
Healthcare
+3
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000