RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Palo Alto, CA seeks an experienced infrastructure engineer to design and optimize large-scale reinforcement learning and post-training workloads. You will collaborate with researchers and engineers to translate ideas into production-grade training pipelines and to improve GPU utilization across the cluster.

This role focuses on building reliable, scalable RL training infrastructures, enhancing monitoring, and ensuring high throughput with multi-node orchestration using modern

Qualifications

  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods).
  • Experience with actor-learner architectures and environment rollout orchestration at scale.
  • Strong Python skills, plus PyTorch or JAX.
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks.
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes).
  • A track record of improving training throughput and GPU utilization at scale.
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases.

Responsibilities

  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production‑grade training pipelines
  • Improve GPU utilization and training throughput across the cluster

Skills

Distributed systems
Actor-learner
Python
PyTorch
JAX
Async training
Replay buffers
Simulation env
Ray
Kubernetes
SLURM
GPU throughput
Debugging

Tools

Ray
SLURM
Kubernetes

Job description

Jobtailor in Palo Alto, CA seeks an experienced infrastructure engineer to design and optimize large-scale reinforcement learning and post-training workloads. You will collaborate with researchers and engineers to translate ideas into production-grade training pipelines and to improve GPU utilization across the cluster.

This role focuses on building reliable, scalable RL training infrastructures, enhancing monitoring, and ensuring high throughput with multi-node orchestration using modern

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Deep Learning Infra Architect for Large-Scale Training
Senior Deep Learning Infra Architect for Large-Scale Training

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000
Reinforcement Learning Infrastructure Engineer
Reinforcement Learning Infrastructure Engineer

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000