RL Scaling Research Scientist — Frontier Models

Speedrun Talent Network

Menlo Park (CA)

On-site

USD 225,000 - 350,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Periodic Labs in Menlo Park (CA) or Montreal seeks a researcher to advance frontier RL for scientific tasks. You will design experiments to understand how RL scales with compute, model size, data, and reward quality, and develop methods from controlled experiments to large-scale runs like Periodic Neon.

This role requires hands-on experience training LLMs with reinforcement learning, a meticulous, scientific approach, and the ability to prototype small-scale RL setups that transfer to bigger

Qualifications

  • Hands-on experience training LLMs with reinforcement learning.
  • Detail-oriented, rigorous scientific approach.
  • Ability to design small-scale RL experiments transferable to large-scale runs.
  • Experience debugging and testing research ideas on a complex training stack.

Responsibilities

  • Design experiments to understand RL scaling with compute, model size, data, and reward quality.
  • Develop RL algorithms across policy optimization, advantage estimation, exploration, and credit assignment.
  • Build adaptive sampling and curriculum methods adjusting task difficulty and rollout counts.
  • Study bias and stability during RL training and address policy staleness and train–inference mismatch.
  • Improve compute efficiency with hyperparameters, length penalties, and update schedules.

Skills

Reinforcement learning
LLM training
Attention to detail
Small-scale RL experiments
Research experimentation

Education

Bachelor's degree

Job description

Periodic Labs in Menlo Park (CA) or Montreal seeks a researcher to advance frontier RL for scientific tasks. You will design experiments to understand how RL scales with compute, model size, data, and reward quality, and develop methods from controlled experiments to large-scale runs like Periodic Neon.

This role requires hands-on experience training LLMs with reinforcement learning, a meticulous, scientific approach, and the ability to prototype small-scale RL setups that transfer to bigger

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Scaling Research Scientist for Frontier Models
RL Scaling Research Scientist for Frontier Models

Periodic Labs • San Francisco (CA)

On-site
USD 225,000 - 350,000
Research Scientist, Scaling RL
Research Scientist, Scaling RL

Periodic Labs • Menlo Park (CA)

On-site
USD 225,000 - 350,000
Research Scientist, Scaling RL
Research Scientist, Scaling RL

Speedrun Talent Network • Menlo Park (CA)

On-site
USD 225,000 - 350,000
Research Scientist, Scaling RL
Research Scientist, Scaling RL

Periodic Labs • San Francisco (CA)

On-site
USD 225,000 - 350,000
Research Engineer - Frontier Models for Scientific Discovery
Research Engineer - Frontier Models for Scientific Discovery

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Frontier RL Research Engineer: Scale & Systems
Frontier RL Research Engineer: Scale & Systems

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Frontier AI Research Engineer — Data & Scaled Training
Frontier AI Research Engineer — Data & Scaled Training

Periodic • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Frontier AI Research Engineer for Scientific Discovery
Frontier AI Research Engineer for Scientific Discovery

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
RL Frontiers Engineer: Scale-Driven Research & Systems
RL Frontiers Engineer: Scale-Driven Research & Systems

Alex Loftus • New York (NY)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Flexible hours
Vacation and parental leave
+1
Frontier ML Research Engineer: Data, Evals & Scaling
Frontier ML Research Engineer: Data, Evals & Scaling

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000