Research Engineer, RL Environments & Training Infra

Preference Model

Seattle (WA)

On-site

USD 200,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
401K match
Lunch provided everyday onsite
Weekly snack orders
Visa sponsorship & relocation support

Job summary

Preference Model is seeking a Research Engineer or Research Scientist located in Seattle, WA, to innovate in the field of self-directed learning for models. The role involves training and evaluating models in custom RL settings and optimizing ML infrastructure.

Ideal candidates will have experience with LLM post-training pipelines and be proficient in Python and modern RL frameworks. Preference Model offers a competitive compensation package, the opportunity to work with top engineers, and a challenging startup environment.

Qualifications

  • Experience running end-to-end LLM post-training pipelines.
  • Proficiency in Python and PyTorch or JAX.
  • Experience with at least one modern RL training framework.
  • Experience building and operating ML infrastructure at scale.

Responsibilities

  • Train and evaluate models on proprietary RL environments.
  • Architect and optimize RL training infrastructure.
  • Design, implement, and test training environments for RL agents.
  • Profile and optimize training runs end-to-end.

Skills

End-to-end LLM post-training pipelines
Python
PyTorch or JAX
Modern RL training framework
ML infrastructure at scale

Job description

Preference Model is seeking a Research Engineer or Research Scientist located in Seattle, WA, to innovate in the field of self-directed learning for models. The role involves training and evaluating models in custom RL settings and optimizing ML infrastructure.

Ideal candidates will have experience with LLM post-training pipelines and be proficient in Python and modern RL frameworks. Preference Model offers a competitive compensation package, the opportunity to work with top engineers, and a challenging startup environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
ML Engineer (New Grad) — RL Environments
ML Engineer (New Grad) — RL Environments

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
New Grad ML Engineer: Design & Build RL Environments
New Grad ML Engineer: Design & Build RL Environments

Preference Model • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+2
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Senior ML Engineer - RL Environments for Frontier LLMs
Senior ML Engineer - RL Environments for Frontier LLMs

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Staff ML Engineer: Low-Level Kernels & RL Environments
Staff ML Engineer: Low-Level Kernels & RL Environments

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research Engineer: Scalable RL Environments & Production
Research Engineer: Scalable RL Environments & Production

somewhere • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Ownership upside
Collaboration with leading AI research organizations
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
RL Researcher - Post-Training for Omni-Model AI
RL Researcher - Post-Training for Omni-Model AI

Nuance Labs • Seattle (WA)

On-site
USD 250,000 - 350,000
Health savings account contributions
15 days of PTO
Lunch, drinks, and snacks provided