RL Infrastructure Engineer: Scale End-to-End ML

Lever, Inc.

Sunnyvale (CA)

On-site

USD 150,000 - 450,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Bonus
401K plan
Paid time off
Parental leave
Employee assistance
Life insurance
Disability insurance

Job summary

Institute of Foundation Models is seeking an RL infrastructure engineer to scale end-to-end RL training systems and extend distributed training frameworks across multi-node, multi-GPU clusters.

You will work alongside researchers and engineers to integrate rollout generation, reward computation, and policy updates, while improving reliability, maintainability, and performance of the training stack.

Qualifications

  • 5+ years of experience in ML systems, infra, or distributed training.
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod).
  • Strong software engineering fundamentals (Python, systems design, testing).
  • Proven multi-node experience (Slurm, Kubernetes, Ray) and debugging skills (NCCL/GLOO).
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs.
  • Experience with ML platform/infrastructure or distributed inference optimization teams.
  • Hands-on experience developing an LLM RL pipeline across rollout/inference and distributed training integration.
  • Working knowledge of policy optimization methods such as PPO or GRPO.

Responsibilities

  • Extend or modify training frameworks to support new use cases and architectures.
  • Connect training workers with rollout generation engines, including trajectory exchange and policy-weight synchronization.
  • Create and debug multi-node launch scripts with flexible batch sizes and hardware targets.
  • Build systems for experiment tracking, job monitoring, and logging for collaborators and researchers.
  • Write production-quality code and tests for ML infra in PyTorch or JAX, ensuring reliability at scale.
  • Implement reward/verifier integration and trajectory processing, validating log probabilities and loss inputs.
  • Coordinate rollout and training workers, including checkpoint/restart and failure recovery; track policy versions and sample staleness.

Skills

ML systems
Distributed training
Python
Systems design
Testing
Multi-node experience
NCCL/GLOO debugging
Policy optimization

Tools

DeepSpeed
FSDP
FairScale
Horovod
Slurm
Kubernetes
Ray

Job description

Institute of Foundation Models is seeking an RL infrastructure engineer to scale end-to-end RL training systems and extend distributed training frameworks across multi-node, multi-GPU clusters.

You will work alongside researchers and engineers to integrate rollout generation, reward computation, and policy updates, while improving reliability, maintainability, and performance of the training stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer — Reinforcement Learning
Machine Learning Engineer — Reinforcement Learning

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical benefits
Dental benefits
Vision benefits
+7
Research Engineer, RL & LLM Post-Training — Scale ML Infra
Research Engineer, RL & LLM Post-Training — Scale ML Infra

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
RL Post-Training Infrastructure Architect
RL Post-Training Infrastructure Architect

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer for Large-Scale Mid-Training & RL
Senior ML Infra Engineer for Large-Scale Mid-Training & RL

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Staff Research Software Engineer: Open ML Infra & RL Training
Staff Research Software Engineer: Open ML Infra & RL Training

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 260,000
Top-tier compensation & equity
Stock options
Health & wellness benefits
+5
RL Post-Training Systems Engineer (Distributed Infra)
RL Post-Training Systems Engineer (Distributed Infra)

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000