Reinforcement Learning Infrastructure Engineer

Jobtailor

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Palo Alto, CA seeks an experienced infrastructure engineer to design and optimize large-scale reinforcement learning and post-training workloads. You will collaborate with researchers and engineers to translate ideas into production-grade training pipelines and to improve GPU utilization across the cluster.

This role focuses on building reliable, scalable RL training infrastructures, enhancing monitoring, and ensuring high throughput with multi-node orchestration using modern

Qualifications

  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods).
  • Experience with actor-learner architectures and environment rollout orchestration at scale.
  • Strong Python skills, plus PyTorch or JAX.
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks.
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes).
  • A track record of improving training throughput and GPU utilization at scale.
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases.

Responsibilities

  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production‑grade training pipelines
  • Improve GPU utilization and training throughput across the cluster

Skills

Distributed systems
Actor-learner
Python
PyTorch
JAX
Async training
Replay buffers
Simulation env
Ray
Kubernetes
SLURM
GPU throughput
Debugging

Tools

Ray
SLURM
Kubernetes

Job description


  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads

  • Improve the reliability, scalability, and throughput of distributed RL training pipelines

  • Build actor-learner architectures and orchestrate environment rollouts at scale

  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems

  • Collaborate with researchers to translate algorithmic ideas into production‑grade training pipelines

  • Improve GPU utilization and training throughput across the cluster


Requirements


  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)

  • Experience with actor-learner architectures and environment rollout orchestration at scale

  • Strong Python skills, plus PyTorch or JAX

  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks

  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)

  • A track record of improving training throughput and GPU utilization at scale

  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
ML Systems Engineer
ML Systems Engineer

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Software Engineer, Code RL
Staff Software Engineer, Code RL

Anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 140,000 - 180,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff - RL Training Framework
Member of Technical Staff - RL Training Framework

Xai • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3