Research Scientist, Reinforcement Learning

Deeproute.ai

Colorado

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deeproute.ai is seeking a skilled professional to advance next-generation autonomous driving systems in Colorado. The candidate will apply reinforcement learning in safety-critical environments, focusing on training RL models and optimizing complex driving behaviors. Key skills include proficiency in modern RL algorithms and Python. Ideal candidates will have publications in top-tier venues and contributions to open-source RL projects. This role offers an opportunity to impact the future of autonomous driving technology significantly.

Qualifications

  • Train and deploy RL policies in closed-loop driving environments.
  • Scale RL training using massively parallel simulation systems.
  • Design and optimize reward functions for complex driving behaviors.

Responsibilities

  • Improve sim-to-real transfer for real-world robustness.
  • Collaborate with cross-functional teams to integrate models into production systems.

Skills

Proficiency in modern RL algorithms: DQN, PPO, SAC, TD3, etc.
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc.
Hands-on experience training reward models and finetuning LLM/VLM/VLA
Knowledge of distributed RL training at scale
Proficiency in Python, comfortable with C++
Proficiency in deep learning frameworks such as PyTorch
Knowledge of sim-to-real transfer techniques and domain randomization
Knowledge of traffic rules, driving behavior modeling

Tools

Ray
Horovod

Job description

We are building next-generation end-to-end autonomous driving systems powered by reinforcement learning.

You will work on applying RL in closed-loop, safety-critical environments, leveraging large-scale simulation and real-world driving data to improve safety, comfort, and robustness.

  • Train and deploy RL policies in closed-loop driving environments
  • Scale RL training using massively parallel simulation systems
  • Design and optimize reward functions for complex driving behaviors
  • Improve sim-to-real transfer for real-world robustness
  • Collaborate with cross-functional teams to integrate models into production systems
Core Technical Skills
  • Proficiency in modern RL algorithms: DQN, PPO, SAC, TD3, etc.
  • Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc.
  • Hands‑on experience training reward models and finetuning LLM/VLM/VLA
  • Knowledge of distributed RL training at scale
  • Proficiency with massively parallel simulation environmentsKnowledge of sim‑to‑real transfer techniques and domain randomization
  • Proficiency in Python, comfortable with C++
  • Proficiency in deep learning frameworks such as PyTorch
  • Experience with distributed training frameworks (Ray, Horovod, etc.)
  • Knowledge of model optimization (quantization, pruning) and CUDA is a plus
  • Knowledge of traffic rules, driving behavior modeling
Preferred Qualifications
  • Publications in top‑tier venues (ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV, ICRA, IROS, etc.)
  • Open‑source contributions to RL libraries or autonomous driving projects
  • Previous experience with LLM fine‑tuning using RLHF
  • Knowledge of safe RL, interpretable AI, or robustness techniques
  • Familiarity with autonomous vehicle regulations and safety standards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Reinforcement Learning
Research Scientist, Reinforcement Learning

DeepRoute • Fremont (CA)

On-site
USD 90,000 - 130,000
Reinforcement learning engineer
Reinforcement learning engineer

Dexmate • United States

On-site
USD 100,000 - 130,000
Autonomous Driving RL Research Scientist
Autonomous Driving RL Research Scientist

DeepRoute • Fremont (CA)

On-site
USD 90,000 - 130,000
Machine Learning Engineer (Reinforcement Learning)
Machine Learning Engineer (Reinforcement Learning)

pony.ai • Colorado

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Paid Time Off (Vacation & Public Holidays)
+1
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

Pony.ai Inc. • Fremont (CA)

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Life Insurance (Basic, Voluntary & AD&D)
+4
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

Albert Invent • Fremont (CA)

On-site
USD 150,000 - 250,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (Traditional and Roth 401k)
Life Insurance (Basic, Voluntary & AD&D)
+4
Algorithm Engineer, Reinforcement Learning
Algorithm Engineer, Reinforcement Learning

Bot Auto • Houston (TX), San Francisco (CA)

On-site
USD 100,000 - 150,000
Comprehensive health insurance
Paid time off
Performance bonuses
+1
Senior Algorithm Engineer, Reinforcement Learning
Senior Algorithm Engineer, Reinforcement Learning

Bot Auto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health insurance
Paid time off
Performance bonuses and equity opportunities
Senior Algorithm Engineer, Reinforcement Learning
Senior Algorithm Engineer, Reinforcement Learning

Botauto • Houston (TX)

On-site
USD 140,000 - 210,000
Health insurance
Equity
Paid time off
Algorithm Engineer, Reinforcement Learning
Algorithm Engineer, Reinforcement Learning

Botauto • Houston (TX)

On-site
USD 140,000 - 210,000
Equity
Healthcare benefits
Performance bonuses