Member of Technical Staff, RL Systems

Goaly

Menlo Park (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Meals and office benefits
Visa sponsorship
Location-based hybrid policy

Job summary

Goaly is hiring for an ambitious role at the intersection of distributed systems, ML infrastructure, and high-performance computing. You will architect scalable RL pipelines, run environment orchestration, and optimize training and inference for long-running experiments in a hybrid, office-based setting in Menlo Park, CA.

You will collaborate with researchers to identify bottlenecks, redesign critical paths, and deliver robust platform capabilities that scale with research demand while

Qualifications

  • A track record building and operating distributed systems, ML infra, or performance-critical backends.

Responsibilities

  • Architect and implement end-to-end RL pipelines coordinating rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.

Skills

Distributed systems
ML infrastructure
Python
C++
Rust
Go
Concurrency
Observability
Profiling
Ownership

Tools

PyTorch
JAX
Kubernetes
Ray
Megatron
DeepSpeed
TensorRT-LLM
vLLM

Job description

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.

About the role

You will build the core platform for agentic reinforcement learning: rollout inference, environment orchestration, distributed training, scheduling, data movement, observability, and recovery. Your goal is to make ambitious RL experiments easy to launch, fast to iterate, efficient to scale, and reliable enough to run for days without constant intervention.

This role sits at the intersection of distributed systems, ML infrastructure, inference, and performance engineering. You will work directly with researchers to find the bottlenecks that matter, then redesign the system so that a local optimization becomes durable leverage for every future run.

What you’ll do
  • Architect and implement end-to-end RL pipelines that coordinate asynchronous rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.

  • Build environment infrastructure for sandboxed and stateful agent workloads, including lifecycle management, isolation, retries, timeouts, replay, checkpoint and restore, and clear distinction between task outcomes and infrastructure failures.

  • Scale high-throughput rollout inference through batching, scheduling, caching, load balancing, disaggregated execution, and efficient model-weight updates.

  • Design resource-management and scheduling systems that place heterogeneous RL workloads efficiently across GPU, CPU, memory, network, and storage constraints.

  • Improve training-inference synchronization, checkpointing, fault recovery, elastic scaling, and long-running job resilience.

  • Profile the full stack and remove bottlenecks in GPU utilization, kernels, communication, serialization, data transfer, storage, and environment throughput.

  • Establish correctness and reproducibility through versioned artifacts, data lineage, idempotent operations, invariant checks, and tests for silent failure modes.

  • Build observability and debugging tools that let researchers answer why a run is slow, unstable, or behaviorally wrong without depending on an infrastructure specialist.

  • Create simple APIs and abstractions that make correct, efficient system use the default while preserving the flexibility needed for fast-moving research.

You may be a good fit if you have
  • A strong record building and operating distributed systems, ML infrastructure, high-performance computing platforms, or performance-critical backend systems.

  • Excellent programming ability in Python plus at least one systems language such as C++, Rust, or Go.

  • Experience reasoning about concurrency, partial failure, backpressure, scheduling, consistency, retries, and observability in production systems.

  • Ability to profile a complex workload, identify the limiting resource, and deliver optimizations that hold under realistic scale and failure conditions.

  • Comfort working across boundaries: research code, model runtimes, infrastructure services, cluster schedulers, and accelerator behavior.

  • Strong ownership and communication, including the ability to turn loosely defined research pain into maintainable platform capabilities.

Strong pluses
  • Experience with distributed RL, large-scale pre-training or post-training, online inference, or asynchronous actor-learner architectures.

  • Familiarity with PyTorch or JAX and systems such as FSDP, Megatron, DeepSpeed, Ray, Kubernetes, vLLM, SGLang, TensorRT-LLM, or similar tools.

  • Knowledge of GPU architecture, CUDA or Triton, NCCL/RCCL, RDMA, InfiniBand, NVLink, or topology-aware scheduling.

  • Experience building sandboxed execution, workflow engines, actor systems, durable runtimes, or multi-agent orchestration.

  • Meaningful contributions to open-source ML systems or infrastructure projects.

How we work
  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

  • Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.

  • Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.

  • Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.

A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.

Equal opportunity

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Member of Technical Staff, Research — Early Career(PHD)
Member of Technical Staff, Research — Early Career(PHD)

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
Member of Technical Staff, New Grad
Member of Technical Staff, New Grad

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Meals and office benefits
Member of Technical Staff, Backend & Product
Member of Technical Staff, Backend & Product

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
Software Engineer, RL Training Infra
Software Engineer, RL Training Infra

OpenAI • California (MO)

Hybrid
USD 180,000 - 260,000
Member of Technical Staff, Cluster Infrastructure
Member of Technical Staff, Cluster Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship