Reinforcement Learning Infrastructure Engineer

Elorian AI

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Elorian AI, an on-site AI research lab in Palo Alto, is seeking an infrastructure engineer to design and build core systems for RL training pipelines. You will own end-to-end training infrastructure, from rollout to observability, partnering with researchers to translate ideas into production-grade pipelines.

You will design scalable RL training, optimize GPU utilization, and develop monitoring tools to ensure reliability.

Qualifications

  • 3+ years of distributed systems experience.
  • Experience with actor-learner architectures and environment rollout orchestration at scale.
  • Strong Python skills, plus PyTorch or JAX.
  • Experience with async training infrastructure, replay buffers, or simulation-based environments.
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes).
  • A track record of improving training throughput and GPU utilization at scale.
  • Strong engineering skills; maintainable code and debugging in large codebases.

Responsibilities

  • Design, build, and optimize the infrastructure powering large-scale RL workloads.
  • Improve reliability, scalability, and throughput of distributed RL training pipelines.
  • Build actor-learner architectures and orchestrate environment rollouts at scale.
  • Develop monitoring and observability tools for high uptime and reproducibility.
  • Collaborate with researchers to translate algorithmic ideas into production pipelines.
  • Improve GPU utilization and training throughput across the cluster.

Skills

Distributed systems
Actor-learner architectures
Python
Code quality
Debug complex codebases

Tools

PyTorch
JAX
Ray
SLURM
Kubernetes

Job description

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.

The Role

We're looking for an infrastructure engineer to design and build the core systems behind how we train our models with reinforcement learning (RL).

You'll own the training infrastructure end to end, from rollout and reward pipelines to orchestration, reliability, and observability. The work spans both the algorithmic side of RL and the systems reality of running distributed training at scale, and you'll partner closely with our research team to keep RL training fast, stable, and dependable for the multimodal, visual reasoning models at the center of our work.

What You Will Do
  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines
  • Improve GPU utilization and training throughput across the cluster
What We're Looking For
  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)
  • Experience with actor-learner architectures and environment rollout orchestration at scale
  • Strong Python skills, plus PyTorch or JAX
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)
  • A track record of improving training throughput and GPU utilization at scale
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases
Preferred qualifications (strong candidates may have some, not all):
  • Experience with multimodal or agentic RL environments
  • Experience with RLHF or reward modeling pipelines
  • A self-directed builder who moves quickly and works across teams in an early-stage setting
Logistics

Location: This role is based on-site in Palo Alto, California.

Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.

Visa sponsorship: We sponsor work visas. We can't promise every case will succeed, but for the right person we'll work through the process with you.

Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
RL Infrastructure Engineer — Frontier AI Research
RL Infrastructure Engineer — Frontier AI Research

Aionia Group • San Francisco (CA)

On-site
USD 300,000 - 500,000
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Member of Technical Staff, Platform Engineering
Member of Technical Staff, Platform Engineering

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 250,000
Healthcare
Relocation support
401k with 4% match
+3
SWE (RL Environments) "Reinforcement Learning"
SWE (RL Environments) "Reinforcement Learning"

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000