RL Infrastructure Engineer for Scalable GPU Training

Elorian AI

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Elorian AI, an on-site AI research lab in Palo Alto, is seeking an infrastructure engineer to design and build core systems for RL training pipelines. You will own end-to-end training infrastructure, from rollout to observability, partnering with researchers to translate ideas into production-grade pipelines.

You will design scalable RL training, optimize GPU utilization, and develop monitoring tools to ensure reliability.

Qualifications

  • 3+ years of distributed systems experience.
  • Experience with actor-learner architectures and environment rollout orchestration at scale.
  • Strong Python skills, plus PyTorch or JAX.
  • Experience with async training infrastructure, replay buffers, or simulation-based environments.
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes).
  • A track record of improving training throughput and GPU utilization at scale.
  • Strong engineering skills; maintainable code and debugging in large codebases.

Responsibilities

  • Design, build, and optimize the infrastructure powering large-scale RL workloads.
  • Improve reliability, scalability, and throughput of distributed RL training pipelines.
  • Build actor-learner architectures and orchestrate environment rollouts at scale.
  • Develop monitoring and observability tools for high uptime and reproducibility.
  • Collaborate with researchers to translate algorithmic ideas into production pipelines.
  • Improve GPU utilization and training throughput across the cluster.

Skills

Distributed systems
Actor-learner architectures
Python
Code quality
Debug complex codebases

Tools

PyTorch
JAX
Ray
SLURM
Kubernetes

Job description

Elorian AI, an on-site AI research lab in Palo Alto, is seeking an infrastructure engineer to design and build core systems for RL training pipelines. You will own end-to-end training infrastructure, from rollout to observability, partnering with researchers to translate ideas into production-grade pipelines.

You will design scalable RL training, optimize GPU utilization, and develop monitoring tools to ensure reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Lead RL Infrastructure Engineer — Scalable GPU Training
Lead RL Infrastructure Engineer — Scalable GPU Training

AMD • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Competitive benefits package
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Staff Engineer, RL Systems & ML Infrastructure
Staff Engineer, RL Systems & ML Infrastructure

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Location-based hybrid policy
Senior RL Infrastructure Engineer - Scalable GPU Systems
Senior RL Infrastructure Engineer - Scalable GPU Systems

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
Reinforcement Learning Infrastructure Engineer
Reinforcement Learning Infrastructure Engineer

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RL Infra Engineer: Scale GPU RL Experiments (Equity)
RL Infra Engineer: Scale GPU RL Experiments (Equity)

Aionia Group • San Francisco (CA)

On-site
USD 300,000 - 500,000
RL Post-Training Infrastructure Architect
RL Post-Training Infrastructure Architect

NVIDIA • Washington

On-site
USD 184,000 - 357,000