Senior RL Infrastructure Engineer - Scalable GPU Systems

Vmax AI Corp

San Francisco (CA)

Hybrid

USD 300,000 - 500,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work arrangement

Job summary

Vmax AI Corp is seeking a strong infrastructure engineer to build the systems layer for RL at scale. You will enable thousands of GPUs to run, debug, and reproduce large-scale RL experiments, tackling training orchestration, data pipelines, and observability.

You will own infra projects end to end—from architecture to deployment—while aligning with ML researchers to translate complex experiments into durable, scalable platforms within our San Francisco office (hybrid option possible).

Qualifications

  • Strong software engineering experience.
  • Experience building infrastructure for LLM inference and RL training.
  • Experience with GPU clusters, distributed training, model serving, or high-throughput inference systems.
  • Familiarity with vLLM, SGLang and modern LLM-RL training frameworks.
  • Strong understanding of system reliability, observability, testing, debugging, and performance optimization.
  • Ability to translate messy experimental workflows into durable infrastructure.
  • Experience building tools, platforms, or services used by other technical users.

Responsibilities

  • Build infrastructure for distributed RL training and inference across thousands of GPUs.
  • Improve the reliability, debuggability, and throughput of RL experiments.
  • Build interfaces that allow researchers and applied ML engineers to launch, inspect, compare, and reproduce experiments easily.
  • Own infrastructure projects end to end, from architecture and implementation through deployment, documentation, and long-term maintenance.
  • Identify and eliminate bottlenecks in training, rollout generation, eval execution, data movement, and cluster utilization.
  • Maintain engineering standards for RL infrastructure, including testing, observability, versioning, and reproducibility.

Skills

Software engineering
LLM infra
GPUs & distributed training
vLLM
SGLang
Reliability & observability
Experiment-driven infra

Job description

Vmax AI Corp is seeking a strong infrastructure engineer to build the systems layer for RL at scale. You will enable thousands of GPUs to run, debug, and reproduce large-scale RL experiments, tackling training orchestration, data pipelines, and observability.

You will own infra projects end to end—from architecture to deployment—while aligning with ML researchers to translate complex experiments into durable, scalable platforms within our San Francisco office (hybrid option possible).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Vmax • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Member of Technical Staff - RL Infrastructure
Member of Technical Staff - RL Infrastructure

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
Member of Technical Staff - RL Infrastructure
Member of Technical Staff - RL Infrastructure

Vmax • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Engineer, RL Systems & ML Infrastructure
Staff Engineer, RL Systems & ML Infrastructure

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Location-based hybrid policy
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Lead RL Infrastructure Engineer — Scalable GPU Training
Lead RL Infrastructure Engineer — Scalable GPU Training

AMD • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Competitive benefits package
RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000