RL Infrastructure Engineer — Frontier AI Research

Aionia Group

San Francisco (CA)

On-site

USD 300,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aionia Group in San Francisco is seeking a Systems Infrastructure Engineer to build scalable infrastructure for RL experiments. This role offers a unique opportunity to work on innovative projects with leading researchers in a well-funded AI company.

The ideal candidate has over 2 years of experience in RL systems, a degree in a related field, and a passion for scalable solutions. Competitive compensation includes a base salary of $300K–$500K plus equity.

Qualifications

  • 2+ years building infrastructure for LLM or RL systems.
  • Experience at a high-engineering-bar organization.
  • Experience with GPU clusters, distributed training, and high-throughput inference systems.

Responsibilities

  • Design and deploy infrastructure for distributed RL training and inference.
  • Improve reliability and throughput for large-scale RL experiments.
  • Establish engineering standards for RL infrastructure.

Skills

Building infrastructure for LLM or RL systems
Hands-on experience with GPU clusters
Familiarity with modern LLM-RL training frameworks
Curiosity and hypothesis-driven thinking

Education

Degree in CS, EECS, Mathematics, or a related field

Tools

vLLM
SGLang
veRL
SkyRL
DeepSpeed

Job description

A rare infrastructure role in a frontier RL research operation.

Compensation: $300K–$500K base + equity. San Francisco, on-site. Hiring urgent.

Opportunity

In a seed‑stage, well‑funded AI company, a small engineering team works with top researchers to automate task objectives and scale learning curriculums across thousands of GPUs.

Own the Systems Layer

This role builds the systems layer that enables researchers and applied ML engineers to run, debug, and reproduce large‑scale RL experiments—including distributed rollouts, training orchestration, inference, evaluation, data pipelines, observability, and reliability.

You will own infrastructure projects end‑to‑end, translating experimental workflows into durable, scalable solutions alongside leading RL minds.

Responsibilities
  • Design and deploy infrastructure for distributed RL training and inference across thousands of GPUs.
  • Improve reliability, debuggability, and throughput for large‑scale RL experiments.
  • Build interfaces for researchers and ML engineers to launch, inspect, compare, and reproduce experiments.
  • Eliminate bottlenecks in training, rollout generation, evaluation, data movement, and cluster utilization.
  • Establish engineering standards for RL infrastructure: testing, observability, versioning, and reproducibility.
Qualifications
  • 2+ years building infrastructure for LLM or RL systems.
  • Experience at a high‑engineering‑bar organization—top AI startups, frontier labs, or Big Tech RL research teams.
  • Hands‑on experience with GPU clusters, distributed training, model serving, or high‑throughput inference systems.
  • Familiarity with vLLM, SGLang, and modern LLM‑RL training frameworks.
  • Degree in CS, EECS, Mathematics, or a related field.
  • Very high level of curiosity and hypothesis‑driven thinking.
Nice to Have
  • Experience collaborating with ML researchers on infrastructure for messy experimental workflows.
  • Evidence of strong independent technical work—open‑source projects, competitions, or notable infrastructure contributions.
  • Familiarity with veRL, SkyRL, Slime, FSDP, DeepSpeed, or similar distributed training frameworks.
Exclusions
  • No hands‑on LLM or RL infrastructure experience.
  • No meaningful GPU or distributed systems exposure.
  • No evidence of strong technical ownership or independent work.
Compensation & Logistics
  • $300,000 – $500,000 base (average offer $400K–$425K)
  • Competitive equity at seed stage.
  • Visa: H1B transfer, OPT, and O‑1 supported.
  • San Francisco, on‑site.
Interview Process
  1. Background, fit, and logistics.
  2. Technical interview—deep dive on infrastructure experience, systems design, and RL/LLM stack.
  3. On‑site loop—1–2 days with researchers and engineering team in San Francisco.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reinforcement Learning Infrastructure Engineer
Reinforcement Learning Infrastructure Engineer

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff, Platform Engineering
Member of Technical Staff, Platform Engineering

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 250,000
Healthcare
Relocation support
401k with 4% match
+3
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000
ML Systems Engineer
ML Systems Engineer

Periodic Labs • Menlo Park (CA)

On-site
USD 300,000 - 400,000
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4