Staff Engineer, Scalable RL Infrastructure

Vmax

San Francisco (CA)

Hybrid

USD 300,000 - 500,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Vmax, an applied research lab, seeks an infrastructure engineer to build the systems layer for large-scale RL. You will enable researchers and ML engineers to run, debug, and reproduce experiments across thousands of GPUs, with a focus on reliability and scalability.

The role emphasizes end-to-end ownership, from architecture to deployment, and collaboration with researchers to translate complex workflows into durable platforms.

Qualifications

  • Strong software engineering skills and ability to translate messy experimental workflows into durable infrastructure.
  • Experience building infrastructure for LLM inference and RL training.
  • Familiarity with GPU clusters, distributed training, model serving, or high-throughput inference systems.
  • Knowledge of vLLM, SGLang and modern LLM-RL training frameworks.
  • Ability to design reliable, observable, well-tested systems for large-scale experiments.

Responsibilities

  • Build infrastructure for distributed RL training and inference across thousands of GPUs.
  • Improve the reliability, debuggability, and throughput of RL experiments.
  • Build interfaces that allow researchers and applied ML engineers to launch, inspect, compare, and reproduce experiments easily.
  • Own infrastructure projects end to end, from architecture and implementation through deployment, documentation, and long-term maintenance.
  • Identify and eliminate bottlenecks in training, rollout generation, eval execution, data movement, and cluster utilization.
  • Maintain engineering standards for RL infrastructure, including testing, observability, versioning, and reproducibility.

Skills

Strong software engineering
Distributed training
Collaboration with researchers
System reliability
Debugging and performance optimization

Tools

vLLM
SGLang
LLM-RL training frameworks
Observability tools

Job description

Vmax, an applied research lab, seeks an infrastructure engineer to build the systems layer for large-scale RL. You will enable researchers and ML engineers to run, debug, and reproduce experiments across thousands of GPUs, with a focus on reliability and scalability.

The role emphasizes end-to-end ownership, from architecture to deployment, and collaboration with researchers to translate complex workflows into durable platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior RL Infrastructure Engineer - Scalable GPU Systems
Senior RL Infrastructure Engineer - Scalable GPU Systems

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
Member of Technical Staff - RL Infrastructure
Member of Technical Staff - RL Infrastructure

Vmax • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Member of Technical Staff - RL Infrastructure
Member of Technical Staff - RL Infrastructure

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, RL Systems & ML Infrastructure
Staff Engineer, RL Systems & ML Infrastructure

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Location-based hybrid policy
RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3