Research Engineer, RL & LLM Post-Training — Scale ML Infra

Preference Model

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
Weekly snack orders
Visa sponsorship & relocation support
Health, vision, dental benefits
401K match

Job summary

Preference Model is building automated ML research engineering to advance RL environments and post-training for large language models. We seek machine learning Research Engineers or Research Scientists to blend research and engineering, implementing novel approaches and shaping research directions.

You will train and evaluate models on proprietary RL environments, architect training infrastructure, and optimize end-to-end runs to accelerate research cycles.

Qualifications

  • Experience running end-to-end LLM post-training pipelines of models sizes 7B+.
  • Proficiency in Python and PyTorch or JAX.
  • Experience with at least one modern RL training framework.
  • Experience building and operating ML infrastructure at scale.

Responsibilities

  • Train and evaluate models on proprietary RL environments to validate data quality and task coverage.
  • Architect and optimize RL training infrastructure and experiment management.
  • Design, implement, and test training environments and methodologies for RL agents.
  • Profile and optimize training runs end-to-end to increase throughput.

Skills

End-to-end LLM post-training
Python
PyTorch or JAX
RL training framework
ML infrastructure

Tools

Verl
OpenRLHF

Job description

Preference Model is building automated ML research engineering to advance RL environments and post-training for large language models. We seek machine learning Research Engineers or Research Scientists to blend research and engineering, implementing novel approaches and shaping research directions.

You will train and evaluate models on proprietary RL environments, architect training infrastructure, and optimize end-to-end runs to accelerate research cycles.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Post-Training ML Research Scientist (RLHF/SFT)
Post-Training ML Research Scientist (RLHF/SFT)

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 181,000 - 226,000
Health coverage
Equity
Generous PTO
+2
LLM Post-Training Research Scientist (SFT & RLHF)
LLM Post-Training Research Scientist (SFT & RLHF)

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Health, dental & vision coverage
Retirement benefits
Learning and development stipend
+2
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Research Engineer, Post-training
Member of Technical Staff - Research Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4
ML Engineer — LLM Post-Training & RL Specialist
ML Engineer — LLM Post-Training & RL Specialist

NewsBreak • Mountain View (CA)

On-site
USD 130,000 - 160,000
Health, dental, and vision care
401(k) plan with company matching
Paid time off and holidays
Senior ML Infrastructure Engineer — Frontier RL & LLM Training
Senior ML Infrastructure Engineer — Frontier RL & LLM Training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
ML Research Engineer — Large-Scale Training & Tools
ML Research Engineer — Large-Scale Training & Tools

The Resume Database • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Staff Research Software Engineer: Open ML Infra & RL Training
Staff Research Software Engineer: Open ML Infra & RL Training

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 260,000
Top-tier compensation & equity
Stock options
Health & wellness benefits
+5
RL Frontiers Engineer: Scale-Driven Research & Systems
RL Frontiers Engineer: Scale-Driven Research & Systems

Alex Loftus • New York (NY)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Flexible hours
Vacation and parental leave
+1
RL Research Scientist - Post-Training on LLMs & Code Models
RL Research Scientist - Post-Training on LLMs & Code Models

AMD • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Benefits at a glance