RL Infra Engineer: Scale GPU RL Experiments (Equity)

Aionia Group

San Francisco (CA)

On-site

USD 300,000 - 500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Aionia Group in San Francisco is seeking a Systems Infrastructure Engineer to build scalable infrastructure for RL experiments. This role offers a unique opportunity to work on innovative projects with leading researchers in a well-funded AI company.

The ideal candidate has over 2 years of experience in RL systems, a degree in a related field, and a passion for scalable solutions. Competitive compensation includes a base salary of $300K–$500K plus equity.

Qualifications

  • 2+ years building infrastructure for LLM or RL systems.
  • Experience at a high-engineering-bar organization.
  • Experience with GPU clusters, distributed training, and high-throughput inference systems.

Responsibilities

  • Design and deploy infrastructure for distributed RL training and inference.
  • Improve reliability and throughput for large-scale RL experiments.
  • Establish engineering standards for RL infrastructure.

Skills

Building infrastructure for LLM or RL systems
Hands-on experience with GPU clusters
Familiarity with modern LLM-RL training frameworks
Curiosity and hypothesis-driven thinking

Education

Degree in CS, EECS, Mathematics, or a related field

Tools

vLLM
SGLang
veRL
SkyRL
DeepSpeed

Job description

Aionia Group in San Francisco is seeking a Systems Infrastructure Engineer to build scalable infrastructure for RL experiments. This role offers a unique opportunity to work on innovative projects with leading researchers in a well-funded AI company.

The ideal candidate has over 2 years of experience in RL systems, a degree in a related field, and a passion for scalable solutions. Competitive compensation includes a base salary of $300K–$500K plus equity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RL Infrastructure Engineer — Frontier AI Research
RL Infrastructure Engineer — Frontier AI Research

Aionia Group • San Francisco (CA)

On-site
USD 300,000 - 500,000
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior RL Infrastructure Engineer - Scalable GPU Systems
Senior RL Infrastructure Engineer - Scalable GPU Systems

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Lead RL Infrastructure Engineer — Scalable GPU Training
Lead RL Infrastructure Engineer — Scalable GPU Training

AMD • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Competitive benefits package
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
AI Systems Engineer — RL Environments & Scalable Infra
AI Systems Engineer — RL Environments & Scalable Infra

AI Talent Now • San Francisco (CA)

On-site
USD 120,000 - 150,000
Infrastructure Engineer - RL Environment Platform
Infrastructure Engineer - RL Environment Platform

bespokelabs • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Competitive salary and equity
Opportunity to work with leading AI research labs