RL Systems Engineer - Scale, Reliability & Observability

United States Digital Space LLC

San Francisco (CA)

Hybrid

USD 500,000 - 850,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

United States Digital Space LLC in San Francisco is seeking an experienced software engineer to design, build, and operate RL-scale distributed systems. You will work across training, sampling, and environment execution on a large fleet, with a focus on reliability and performance.

You will collaborate with researchers and performance engineers to preserve training correctness, reduce tail latency, and implement observability, fault tolerance, and automation across the stack.

Qualifications

  • Strong software engineering skills in Python and at least one systems language (Rust, C++, or Go).
  • Experience designing, building, and operating large-scale distributed systems in production.
  • Deep understanding of distributed systems fundamentals, including consistency, coordination, and failure modes.
  • Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and network.
  • Experience debugging complex failures across many hosts and services.
  • Strong written communication, including design documents and incident writeups.

Responsibilities

  • Design, build, and operate the distributed systems that run RL at scale, across training, sampling, and environment execution
  • Find and remove whatever currently limits the system, whether it's scheduling, data movement, storage, networking, or coordination
  • Build fault tolerance into every layer: failure detection, isolation, and recovery that keep long-running jobs making progress without human intervention
  • Design resource management and autoscaling so that compute follows demand as a run's needs shift
  • Build observability that makes it possible to understand what a run is doing and why it slowed down, stalled, or produced unexpected results
  • Build automation that detects and remediates common problems, and design interfaces that let engineers and automated tools operate runs safely
  • Work with researchers and performance engineers to make sure systems changes preserve training correctness and don't introduce subtle nondeterminism
  • Remove classes of failure at their source through incident review, testing, and redesign, and write clear design documents for what you build

Skills

Python
Rust
C++
Go
Distributed systems
Performance analysis
Debugging
Technical writing

Tools

Kubernetes

Job description

United States Digital Space LLC in San Francisco is seeking an experienced software engineer to design, build, and operate RL-scale distributed systems. You will work across training, sampling, and environment execution on a large fleet, with a focus on reliability and performance.

You will collaborate with researchers and performance engineers to preserve training correctness, reduce tail latency, and implement observability, fault tolerance, and automation across the stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Frontier RL Research Engineer: Scale & Systems
Frontier RL Research Engineer: Scale & Systems

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Scale-Out RL Systems Engineer
Scale-Out RL Systems Engineer

Anthropic • New York (NY)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching (optional)
Generous vacation and parental leave
+1
Staff RL Environments Engineer
Staff RL Environments Engineer

Scale AI, Inc. • San Francisco (CA)

On-site
USD 252,000 - 315,000
Health coverage
Dental coverage
Vision coverage
+4
RL Systems Engineer - Scale & Post-Training
RL Systems Engineer - Scale & Post-Training

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000
Senior Site Reliability Engineer for Global, Scalable Systems
Senior Site Reliability Engineer for Global, Scalable Systems

United States Digital Space LLC • New York (NY)

On-site
USD 183,000 - 247,000
RL Environment Infrastructure Engineer — Sandbox & Scale
RL Environment Infrastructure Engineer — Sandbox & Scale

Commergence • Colorado

Hybrid
USD 150,000 - 210,000
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
RL Systems Engineer: Environments & High-Throughput Pipelines
RL Systems Engineer: Environments & High-Throughput Pipelines

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Staff RL Environments Engineer: End-to-End Systems
Staff RL Environments Engineer: End-to-End Systems

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 252,000 - 315,000
Health insurance
Dental & vision coverage
:Learning stipend
RL Research Engineer: Scalable Training & Eval (Remote)
RL Research Engineer: Scalable Training & Eval (Remote)

24-MAG • United States

Remote
USD 400,000 - 800,000