RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

NVIDIA seeks an engineer to architect and scale RL post-training infrastructure from experimentation on a single GPU to production across thousands of nodes. You will optimize training-inference loops on GPUs, CPUs, and LPUs while improving open-source RL frameworks and partnering with cross-team groups.

The role requires expertise in distributed systems, Python and C/C++, and experience delivering production-grade runtime frameworks at large AI labs or hyperscalers. Hybrid work is supported.

Qualifications

  • MS or PhD in Computer Science, Computer Engineering, or a related field (or equivalent experience)
  • 5+ years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering
  • Strong proficiency in Python and C/C++
  • Demonstrated experience building or contributing to large-scale distributed systems or runtime frameworks in production at a frontier AI lab, hyperscaler, or major technology company
  • Strong verbal and written communication skills and the ability to collaborate across organizational and geographic boundaries

Responsibilities

  • Architect and build RL post-training infrastructure that scales from a single GPU to thousands of nodes
  • Tune RL training-inference-rollout loops for performance on GPUs, CPUs, and LPUs
  • Improve performance and usability of open-source RL frameworks
  • Collaborate with NVIDIA teams (networking, math libraries, compilers) and hardware teams to enable next-gen post-training workloads

Skills

Python
C/C++
Distributed systems
Communication

Education

MS or PhD in CS/CE

Tools

Kubernetes
PyTorch
Ray
NCCL

Job description

NVIDIA seeks an engineer to architect and scale RL post-training infrastructure from experimentation on a single GPU to production across thousands of nodes. You will optimize training-inference loops on GPUs, CPUs, and LPUs while improving open-source RL frameworks and partnering with cross-team groups.

The role requires expertise in distributed systems, Python and C/C++, and experience delivering production-grade runtime frameworks at large AI labs or hyperscalers. Hybrid work is supported.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Post-Training Systems Engineer (Distributed Infra)
RL Post-Training Systems Engineer (Distributed Infra)

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
RL Post-Training Infrastructure Architect
RL Post-Training Infrastructure Architect

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Senior RL Post-Training Frameworks Architect
Senior RL Post-Training Frameworks Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Washington

On-site
USD 184,000 - 357,000