Research Scientist - Reinforcement Learning

Optimized, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 185,000 - 255,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Optimized, Inc. is building a small, talent-dense team in San Francisco to deploy AI agents into critical supply chains. You will own RL and post-training, shaping reward models, training loops, and evaluations that translate model capability into reliable long-horizon decisions.

This role is on-site 5 days a week in San Francisco, with visa transfers supported. Compensation ranges from $185,000 to $255,000 plus competitive equity, and you will ground learning in real deployment data and ship it

Qualifications

  • Have a PhD or equivalent research experience in RL, ML, or a related field.
  • Have hands-on experience with reinforcement learning, post-training, or RLHF for LLMs.
  • Are comfortable building research prototypes in Python and iterating quickly.
  • Understand reward modeling, policy optimization, and evaluation of sequential decision-making.
  • Care about real-world impact, and you have driven research through to production.
  • Are excited about applying AI to complex, messy, real-world optimization problems.

Responsibilities

  • Train agents to act: design and run RL and post-training pipelines that improve how agents plan and execute multi-step work.
  • Build reward models: define and train the reward signals that capture what a good supply chain decision looks like.
  • Evaluate long-horizon behavior: build evals that measure agent reliability across long, high-stakes workflows.
  • Ground learning in reality: use real deployment data and feedback to close the gap between simulation and production.
  • Ship research to production: work with engineers to bring training breakthroughs into the live agent platform.

Skills

Reinforcement learning
Python
RLHF
Policy optimization
Sequential decision-making

Education

PhD or equivalent RL/ML research

Tools

PyTorch
TensorFlow

Job description

We're building a small, talent-dense team to deploy AI agents into the most critical supply chains in the world.

You'll own RL and post-training: the reward models, training loops, and evaluations that turn raw model capability into reliable long-horizon decision-making. You'll ground that work in real deployment data and ship it into production.

This role is on-site 5 days a week in San Francisco. Visa transfers (OPT, H-1B) are supported. Compensation range is $185,000 - $255,000 + competitive equity.

What you'll do
  • Train agents to act: You'll design and run RL and post-training pipelines that improve how the agents plan and execute multi-step work.
  • Build reward models: You'll define and train the reward signals that capture what a good supply chain decision actually looks like.
  • Evaluate long-horizon behavior: You'll build evals that measure agent reliability across long, high-stakes workflows, not just single turns.
  • Ground learning in reality: You'll use real deployment data and feedback to close the gap between simulation and production.
  • Ship research to production: You'll work with engineers to bring training breakthroughs into the live agent platform.
What you'll bring
  • Have a PhD or equivalent research experience in RL, ML, or a related field
  • Have hands-on experience with reinforcement learning, post-training, or RLHF for LLMs
  • Are comfortable building research prototypes in Python and iterating quickly
  • Understand reward modeling, policy optimization, and evaluation of sequential decision-making
  • Care about real-world impact, and you have driven research through to production
  • Are excited about applying AI to complex, messy, real-world optimization problems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production RL Research Scientist: Supply Chains
Production RL Research Scientist: Supply Chains

Optimized, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 185,000 - 255,000
Research Engineer - Agents
Research Engineer - Agents

Optimized, Inc. • San Francisco (CA)

On-site
USD 175,000 - 240,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Engineer
Software Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 225,000 - 275,000
Equity
RL Researcher, Product Systems
RL Researcher, Product Systems

Autohand AI Ltd. • San Francisco (CA)

On-site
USD 150,000 - 210,000
Publications
Research Scientist, LLM Evaluation & Post-Training
Research Scientist, LLM Evaluation & Post-Training

OneForma • United States

Hybrid
USD 140,000 - 210,000
Research Scientist - Agent Systems
Research Scientist - Agent Systems

Optimized, Inc. • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Scientist - Long-Horizon Multi-Agent Systems
Research Scientist - Long-Horizon Multi-Agent Systems

techire ai • San Francisco (CA)

On-site
USD 400,000 - 450,000
Staff MLE, Reinforcement Learning
Staff MLE, Reinforcement Learning

People In AI • San Francisco (CA)

Hybrid
USD 270,000 - 280,000
AIML - Machine Learning Research Lead, RL Agents, MLR
AIML - Machine Learning Research Lead, RL Agents, MLR

Socket.dev • Cupertino (CA)

On-site
USD 220,000 - 320,000