RL Infrastructure Engineer — Scale & Production

Inception

San Francisco (CA)

On-site

USD 200,000 - 350,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision insurance
Catered meals (breakfast, lunch, &
Equity
Flexible vacation and PTO

Job summary

Inception is seeking engineers and scientists to design, optimize, and maintain core systems enabling scalable, efficient reinforcement learning for large models.

This role sits at the intersection of research and large-scale systems engineering, with duties ranging from optimizing rollout and reward pipelines to improving reliability, observability, and orchestration in close collaboration with researchers.

Qualifications

  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • Understanding of ML frameworks from a systems perspective (PyTorch, TensorFlow, Ray, Megatron).
  • Experience with reinforcement learning workloads (PPO, DPO, RLHF, or reward modeling).
  • Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.

Responsibilities

  • Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
  • Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
  • Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.

Skills

ML systems understanding
Reinforcement learning workloads
Distributed systems
Experience with RL pipelines
CI/CD
PyTorch/TensorFlow
Ray/Megatron

Education

BS/MS/PhD in CS or Engineering (or equivalent)

Tools

Docker
Kubernetes
CI/CD tooling
Ray
Megatron
PyTorch
TensorFlow

Job description

Inception is seeking engineers and scientists to design, optimize, and maintain core systems enabling scalable, efficient reinforcement learning for large models.

This role sits at the intersection of research and large-scale systems engineering, with duties ranging from optimizing rollout and reward pipelines to improving reliability, observability, and orchestration in close collaboration with researchers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer — Scale Research to Production
ML Infrastructure Engineer — Scale Research to Production

Jobzhr • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 500,000
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
RL Training Infra Engineer: Scale, Debug, & Optimize
RL Training Infra Engineer: Scale, Debug, & Optimize

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
Infrastructure Research Engineer - Large-Scale RL Systems
Infrastructure Research Engineer - Large-Scale RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000