Research Scientist, Reinforcement Learning

DeepRoute

Fremont (CA)

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DeepRoute in Fremont, California is seeking engineers to develop end-to-end autonomous driving systems using reinforcement learning. The ideal candidate will have experience training RL policies in safety-critical environments and expertise in modern RL algorithms like DQN and PPO. You will collaborate with teams to optimize driving behaviors, build parallel simulations, and improve safety. Knowledge of deep learning frameworks like PyTorch and experience with LLM fine-tuning are beneficial.

Qualifications

  • Experience with massively parallel simulation environments.
  • Knowledge of sim-to-real transfer techniques and domain randomization.
  • Knowledge of model optimization (quantization, pruning) and CUDA is a plus.

Responsibilities

  • Train and deploy RL policies in closed-loop driving environments.
  • Scale RL training using massively parallel simulation systems.
  • Design and optimize reward functions for complex driving behaviors.
  • Improve sim-to-real transfer for real-world robustness.
  • Collaborate with cross-functional teams to integrate models into production systems.

Skills

Proficiency in modern RL algorithms: DQN, PPO, SAC, TD3
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO
Hands-on experience training reward models
Knowledge of distributed RL training at scale
Proficiency in Python, comfortable with C++
Proficiency in deep learning frameworks such as PyTorch

Tools

Ray
Horovod

Job description

We are building next-generation end-to-end autonomous driving systems powered by reinforcement learning.

You will work on applying RL in closed-loop, safety-critical environments, leveraging large-scale simulation and real-world driving data to improve safety, comfort, and robustness.

  • Train and deploy RL policies in closed-loop driving environments
  • Scale RL training using massively parallel simulation systems
  • Design and optimize reward functions for complex driving behaviors
  • Improve sim-to-real transfer for real-world robustness
  • Collaborate with cross-functional teams to integrate models into production systems
Core Technical Skills
  • Proficiency in modern RL algorithms: DQN, PPO, SAC, TD3, etc.
  • Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc.
  • Hands‑on experience training reward models and finetuning LLM/VLM/VLA
  • Knowledge of distributed RL training at scale
  • Proficiency with massively parallel simulation environmentsKnowledge of sim‑to‑real transfer techniques and domain randomization
  • Proficiency in Python, comfortable with C++
  • Proficiency in deep learning frameworks such as PyTorch
  • Experience with distributed training frameworks (Ray, Horovod, etc.)
  • Knowledge of model optimization (quantization, pruning) and CUDA is a plus
  • Knowledge of traffic rules, driving behavior modeling
Preferred Qualifications
  • Publications in top‑tier venues (ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV, ICRA, IROS, etc.)
  • Open‑source contributions to RL libraries or autonomous driving projects
  • Previous experience with LLM fine‑tuning using RLHF
  • Knowledge of safe RL, interpretable AI, or robustness techniques
  • Familiarity with autonomous vehicle regulations and safety standards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Reinforcement Learning
Research Scientist, Reinforcement Learning

Deeproute.ai • Colorado

On-site
USD 100,000 - 150,000
Reinforcement learning engineer
Reinforcement learning engineer

Dexmate • United States

On-site
USD 100,000 - 130,000
Autonomous Driving RL Research Scientist
Autonomous Driving RL Research Scientist

DeepRoute • Fremont (CA)

On-site
USD 90,000 - 130,000
Machine Learning Engineer (Reinforcement Learning)
Machine Learning Engineer (Reinforcement Learning)

pony.ai • Colorado

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Paid Time Off (Vacation & Public Holidays)
+1
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

Pony.ai Inc. • Fremont (CA)

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Life Insurance (Basic, Voluntary & AD&D)
+4
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

Albert Invent • Fremont (CA)

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Life Insurance (Basic, Voluntary & AD&D)
+4
Senior Algorithm Engineer, Reinforcement Learning
Senior Algorithm Engineer, Reinforcement Learning

Bot Auto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health insurance
Paid time off
Performance bonuses and equity opportunities
Algorithm Engineer, Reinforcement Learning
Algorithm Engineer, Reinforcement Learning

Bot Auto • Houston (TX), San Francisco (CA)

On-site
USD 100,000 - 150,000
Comprehensive health insurance
Paid time off
Performance bonuses
+1
RL Research Engineer, Self-Driving Autonomy & Robotics
RL Research Engineer, Self-Driving Autonomy & Robotics

Applied Intuition Inc. • Sunnyvale (CA)

Hybrid
USD 126,000 - 423,000
Research Engineer / Scientist – Reinforcement Learning (RL)
Research Engineer / Scientist – Reinforcement Learning (RL)

Percepta • New York (NY)

On-site
USD 110,000 - 150,000