Staff Engineer, RL Systems & ML Infrastructure

Goaly

Menlo Park (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Meals and office benefits
Visa sponsorship
Location-based hybrid policy

Job summary

Goaly is hiring for an ambitious role at the intersection of distributed systems, ML infrastructure, and high-performance computing. You will architect scalable RL pipelines, run environment orchestration, and optimize training and inference for long-running experiments in a hybrid, office-based setting in Menlo Park, CA.

You will collaborate with researchers to identify bottlenecks, redesign critical paths, and deliver robust platform capabilities that scale with research demand while

Qualifications

  • A track record building and operating distributed systems, ML infra, or performance-critical backends.

Responsibilities

  • Architect and implement end-to-end RL pipelines coordinating rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.

Skills

Distributed systems
ML infrastructure
Python
C++
Rust
Go
Concurrency
Observability
Profiling
Ownership

Tools

PyTorch
JAX
Kubernetes
Ray
Megatron
DeepSpeed
TensorRT-LLM
vLLM

Job description

Goaly is hiring for an ambitious role at the intersection of distributed systems, ML infrastructure, and high-performance computing. You will architect scalable RL pipelines, run environment orchestration, and optimize training and inference for long-running experiments in a hybrid, office-based setting in Menlo Park, CA.

You will collaborate with researchers to identify bottlenecks, redesign critical paths, and deliver robust platform capabilities that scale with research demand while

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Senior RL Infrastructure Engineer - Scalable GPU Systems
Senior RL Infrastructure Engineer - Scalable GPU Systems

Vmax AI Corp • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Hybrid work arrangement
Research Scientist, Agentic RL & Scalable AI
Research Scientist, Agentic RL & Scalable AI

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
RL Environment Platform Engineer
RL Environment Platform Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 100,000 - 140,000
Health coverage
Opportunity to work with leading AI research labs
ML Systems Engineer — On-Site in Palo Alto, High-Impact
ML Systems Engineer — On-Site in Palo Alto, High-Impact

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers