Staff Engineer, Scalable RL Infrastructure

Inception

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inception is seeking engineers and scientists to design, optimize, and maintain core systems enabling scalable reinforcement learning for large models. You will work at the intersection of research and systems engineering, optimizing rollout and reward pipelines, improving reliability, observability, and orchestration, and ensuring production-readiness.

Responsibilities include building infrastructure for RL workloads, boosting training throughput, and creating shared tooling for monitoring and

Qualifications

  • Understanding of RL systems from a systems perspective.
  • Experience with large-scale ML workflows and orchestration.
  • Familiarity with reliability, observability, and production readiness.

Responsibilities

  • Design, build, and optimize infrastructure for large-scale RL workloads.
  • Improve reliability and throughput of RL training pipelines.
  • Develop monitoring and observability tools for high uptime and debuggability.

Skills

RL systems
Distributed systems
Observability
Orchestration

Education

BS/MS/PhD in CS/Engineering

Tools

Docker
Kubernetes
CI/CD

Job description

Inception is seeking engineers and scientists to design, optimize, and maintain core systems enabling scalable reinforcement learning for large models. You will work at the intersection of research and systems engineering, optimizing rollout and reward pipelines, improving reliability, observability, and orchestration, and ensuring production-readiness.

Responsibilities include building infrastructure for RL workloads, boosting training throughput, and creating shared tooling for monitoring and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Research Software Engineer: Scalable RL Training Systems
Research Software Engineer: Scalable RL Training Systems

Jobzhr • New York (NY)

On-site
USD 180,000 - 240,000
Salary and equity
Stock options
Healthcare
+3
Infrastructure Research Engineer - Large-Scale RL Systems
Infrastructure Research Engineer - Large-Scale RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Staff Software Engineer — RL Infrastructure & Platforms
Staff Software Engineer — RL Infrastructure & Platforms

Anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 140,000 - 180,000