Machine Learning Engineer — Reinforcement Learning

Lever, Inc.

Sunnyvale (CA)

On-site

USD 150,000 - 450,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Bonus
401K plan
Paid time off
Parental leave
Employee assistance
Life insurance
Disability insurance

Job summary

Institute of Foundation Models is seeking an RL infrastructure engineer to scale end-to-end RL training systems and extend distributed training frameworks across multi-node, multi-GPU clusters.

You will work alongside researchers and engineers to integrate rollout generation, reward computation, and policy updates, while improving reliability, maintainability, and performance of the training stack.

Qualifications

  • 5+ years of experience in ML systems, infra, or distributed training.
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod).
  • Strong software engineering fundamentals (Python, systems design, testing).
  • Proven multi-node experience (Slurm, Kubernetes, Ray) and debugging skills (NCCL/GLOO).
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs.
  • Experience with ML platform/infrastructure or distributed inference optimization teams.
  • Hands-on experience developing an LLM RL pipeline across rollout/inference and distributed training integration.
  • Working knowledge of policy optimization methods such as PPO or GRPO.

Responsibilities

  • Extend or modify training frameworks to support new use cases and architectures.
  • Connect training workers with rollout generation engines, including trajectory exchange and policy-weight synchronization.
  • Create and debug multi-node launch scripts with flexible batch sizes and hardware targets.
  • Build systems for experiment tracking, job monitoring, and logging for collaborators and researchers.
  • Write production-quality code and tests for ML infra in PyTorch or JAX, ensuring reliability at scale.
  • Implement reward/verifier integration and trajectory processing, validating log probabilities and loss inputs.
  • Coordinate rollout and training workers, including checkpoint/restart and failure recovery; track policy versions and sample staleness.

Skills

ML systems
Distributed training
Python
Systems design
Testing
Multi-node experience
NCCL/GLOO debugging
Policy optimization

Tools

DeepSpeed
FSDP
FairScale
Horovod
Slurm
Kubernetes
Ray

Job description

About the Institute of Foundation Models

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries.Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub forhigh-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AIpioneers.

The Role

We’re looking for an RL infrastructure engineer to help extend and scale our end-to-end RL training systems. You’ll work side-by-side with world-class researchers and engineers to:

  • Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
  • Integrate rollout generation, reward computation, trajectory processing, policy updates, and weight synchronization
  • Build robust config + launch systems across multi-node, multi-GPU clusters
  • Own experiment tracking, metrics logging, and job monitoring for external visibility
  • Improve training system reliability, maintainability, and performance

Hands-on experience developing end-to-end RL infrastructure is required. Strong infrastructure and systems experience is what we value most.

Key Responsibilities
  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
  • Training & Inference Integration – Connect training workers with rollout generation engines (e.g., vLLM, SGLang), including trajectory exchange and policy-weight synchronization.
  • Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
  • Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
  • RL Pipeline Development – Implement reward/verifier integration and trajectory processing, and validate log probabilities, token masks, and loss inputs with researchers.
  • RL Execution & Recovery – Coordinate rollout and training workers, including checkpoint/restart and failure recovery; track policy versions and sample staleness when using asynchronous execution.
Qualifications
Must-Haves:
  • 5+ years of experience in ML systems, infra, or distributed training
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
  • Strong software engineering fundamentals (Python, systems design, testing)
  • Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs
  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
  • Experience with large-scale machine learning workloads (strong ML fundamentals)
  • Hands-on experience developing an LLM RL pipeline, with ownership across rollout/inference and distributed training integration.
  • Working knowledge of policy optimization methods such as PPO or GRPO, sufficient to implement and debug sampling, log probabilities, and policy updates.
Nice-to-Haves:
  • Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
  • Familiarity with performance profiling, kernel fusion, or memory optimization
  • Open-source contributions or published research (MLSys, ICML, NeurIPS)
  • CUDA or Triton kernel experience
  • Experience with large-scale pre-training
  • Experience building custom training pipelines at scale and modifying them for custom needs
  • Deep familiarity with training infrastructure and performance tuning
  • Experience with asynchronous RL, multi-turn or tool-using rollout environments, or custom reward/verifier systems.

$150,000 - $450,000 a year

Salary Range

The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.

The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.

Benefits Include

*Comprehensive medical, dental, and vision benefits

*Bonus

*401K Plan

*Generous paid time off, sick leave and holidays

*Paid Parental Leave

*Employee Assistance Program

*Life insurance and disability

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Research Engineer / Research Scientist, RL Frontiers
Research Engineer / Research Scientist, RL Frontiers

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Research Scientist - Reinforcement Learning
Research Scientist - Reinforcement Learning

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Research, RL Scaling
Research, RL Scaling

Thinking Machines Lab Inc. • San Francisco (CA)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites
Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 300,000 - 600,000
Comprehensive benefits
Bonus
401K Plan
+4
Research Engineer / Performance Engineer, RL Distributed Systems
Research Engineer / Performance Engineer, RL Distributed Systems

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Research Engineer, Post-training
Member of Technical Staff - Research Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4