Founding Reinforcement Learning Engineer

Wwshemi

San Francisco (CA)

On-site

USD 125,000 - 200,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Wwshemi in San Francisco, CA seeks a founding engineer to build core reinforcement learning systems from the ground up. You will own environment design, model training, evaluation, and help shape the company's technical direction.

You will work directly with the founders, turning research ideas into production, and scaling training on GPU clusters while refining evaluation frameworks and contributing to engineering culture as the team grows.

Qualifications

  • Hands-on RL/ML engineering experience (2+ years).
  • Experience with policy gradient methods and RLHF is a plus.
  • Strong Python and ML framework skills expected.
  • Familiarity with distributed training and GPU infra.

Responsibilities

  • Design and build RL environments, reward functions, and training pipelines.
  • Train and fine-tune models using PPO, GRPO, DPO, RLHF, and RLAIF.
  • Develop evaluation frameworks to measure agent performance.
  • Scale training on GPU clusters and maintain pipelines.
  • Turn research ideas into production systems.

Skills

Python
PyTorch
JAX
Reinforcement learning
RLHF
Ray
CUDA
Kubernetes
Gymnasium
RLlib
MuJoCo
OpenRLHF
Distributed training

Education

CS/Math/Physics degree
Master/PhD preferred

Tools

Ray
CUDA
Kubernetes

Job description

About the Role

Build core reinforcement learning systems from the ground up as a founding engineer on an early-stage AI team. You will work directly with the founders and own the process from environment design and model training through evaluation, helping shape the company's technical direction.

What You'll Do
  • Design and build reinforcement learning environments, reward functions, and training pipelines.
  • Train and fine-tune models using methods such as PPO, GRPO, DPO, RLHF, and RLAIF.
  • Develop evaluation frameworks to measure model and agent performance.
  • Run experiments, interpret results, and decide which approaches to explore next.
  • Scale training on GPU clusters and maintain reliable pipelines.
  • Turn research ideas into production systems.
  • Help establish engineering culture and hire future engineers.
What We're Looking For
  • At least 2 years of hands-on experience in reinforcement learning or machine learning engineering, with relevant experience potentially ranging from 2 to 10 or more years.
  • Strong Python skills and deep experience with PyTorch or JAX.
  • Practical experience training models with reinforcement learning, including policy gradient methods, reward modeling, or RLHF.
  • Comfort with distributed training and GPU infrastructure, including tools such as Ray, CUDA, and Kubernetes.
  • Experience with LLM post-training or agent training is a plus, as is familiarity with Gymnasium, Ray RLlib, Isaac, MuJoCo, TRL, verl, or OpenRLHF.
  • A degree in computer science, mathematics, physics, or a related field is sought; a master's or PhD is a plus. Publications or open-source work in RL are also valued.
  • A hands-on builder who enjoys moving quickly in a small, early-stage team.
Compensation & Benefits

Compensation is $125,000 to $200,000 USD annually.

Location

On-site in San Francisco, California, United States.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Reinforcement Learning Engineer
Founding Reinforcement Learning Engineer

Clera • San Francisco (CA)

On-site
USD 125,000 - 200,000
Founding Reinforcement Learning Engineer Clera · San Francisco, CA Full-time · On-site $125,000–200,000 3 hours ago
Founding Reinforcement Learning Engineer Clera · San Francisco, CA Full-time · On-site $125,000–200,000 3 hours ago

Emploive • San Francisco (CA)

On-site
USD 125,000 - 200,000
Founding RL Engineer
Founding RL Engineer

Clera Labs, Inc. • San Francisco (CA)

On-site
USD 125,000 - 200,000
Founding RL Engineer – Build Core RL Systems (SF)
Founding RL Engineer – Build Core RL Systems (SF)

Clera Labs, Inc. • San Francisco (CA)

On-site
USD 125,000 - 200,000
Software Engineer, RL Environments
Software Engineer, RL Environments

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 260,000 - 290,000
Full-Stack Software Engineer, Reinforcement Learning (Mid-Level)
Full-Stack Software Engineer, Reinforcement Learning (Mid-Level)

Clera • San Francisco (CA)

On-site
USD 130,000 - 225,000
Health coverage
Paid time off
Retirement benefits
+1
Machine Learning Engineer — Reinforcement Learning
Machine Learning Engineer — Reinforcement Learning

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical benefits
Dental benefits
Vision benefits
+7
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

On-site
USD 180,000 - 220,000
RL Environment Software Engineer
RL Environment Software Engineer

TalentPluto, Inc. • San Francisco (CA)

Remote
USD 180,000 - 220,000
Research Engineer, Training and Environment Infrastructure
Research Engineer, Training and Environment Infrastructure

Ersilia • San Francisco (CA)

On-site
USD 150,000 - 350,000
Meaningful equity grants
Health, dental, and vision coverage