Reinforcement Learning Environment Engineer

Open Data Science

San Francisco (CA)

Remote

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A company specializing in AI training data is seeking a Reinforcement Learning Environment Engineer to design and build MLE/SWE environments. This remote contractor position requires strong Python skills, hands-on LLM experience, and the ability to operate independently. With an expected deliverable of one task every 8-10 hours, candidates need to have advanced English proficiency and a minimum of 4 hours overlap with Pacific time. There's potential for conversion to full-time employment based on performance.

Qualifications

  • Proven experience with engineering-quality Python.
  • History of shipping and operating production-level LLM/GenAI systems.
  • Ability to meet throughput expectations.

Responsibilities

  • Design and build MLE/SWE environments.
  • Target specified language models and ensure difficulty distribution.
  • Deliver approximately one task every 8-10 hours after onboarding.
  • Edit tasks based on customer feedback within 24 hours.

Skills

Strong Python
Hands-on LLM/GenAI experience
End-to-end pipeline ownership
Advanced English (C1/C2)

Job description

Reinforcement Learning Environment Engineer

RL Environments; MLE; LLM Tasks; Difficulty Distribution; Remote Contractor; PST Overlap (≥4h); Advanced English (C1/C2);

We’re hiring RL Environments Engineers to design and build MLE/SWE environments that deliver high-quality, diverse tasks with minimal supervision. You will target a specific language model, meet a defined difficulty distribution, and deliver about one task every 10 hours. This is a remote contractor role with ≥4 hours overlap to PST and advanced English (C1/C2) required.

About the company

Preference Model is building the next generation of training data to power the future of AI. Today's models are powerful but fail to reach their potential across diverse use cases because so many of the tasks that we want to use these models for are outside of their training data distribution. Preference Model creates reinforcement learning environments that encapsulate real-world use cases, enabling AI systems to practice, adapt, and learn from feedback grounded in reality. We seek to bring the real world into distribution for the models.

Our founding team has previous experience on Anthropic’s data team building data infrastructure, tokenizers, and datasets behind the Claude model. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.

The company is backed by Tier 1 Silicon Valley VC.

Responsibilities
  • Design and build MLE/SWE environments and diverse tasks.
  • Target a specified language model and satisfy the required difficulty distribution.
  • Deliver ~1 task per 8-10 hours once onboarded.
  • Edit tasks within 24 hours based on customer feedback.
  • Onboard quickly and start delivering on day one with minimal supervision.
Requirements
What we’re looking for (must-haves)
  • Strong Python (engineering-quality, not notebook‑only).
  • Hands‑on LLM/GenAI work in production: you’ve shipped and operated real systems (not “wrapped an API and called it AI”).
  • Strong product/engineering ownership: comfortable building, fixing, and scaling end‑to‑end pipelines.
  • ≥4 hours PST overlap and advanced English (C1/C2) for specs, reviews, and feedback.
  • Ability to meet throughput expectations and respond quickly to feedback.
Strong signals (nice‑to‑have, big plus)
  • Experience in high‑stakes or regulated domains (e.g., healthcare, finance, fraud/risk, safety‑critical systems).
  • Experience designing environments/tasks for RL and/or evaluations.
  • Exposure to RL / bandits / agentic systems (not required, but a strong signal).
Not a fit if
  • You’re primarily a prompt engineer without strong ML/engineering foundations.
  • You’re a research‑only / academic‑only profile with little or no shipping/production ownership.
  • You’ve only built in notebooks or rely heavily on managed AutoML tools.
Working conditions
  • hours/week - full time - need 4 hours overlap in the working hours with the team in Pacific time zone;
  • Deliverables-driven; begin shipping on day one.
  • Conversion & relocation: Potential path to FTE and relocation to the Bay Area if performance and mutual fit align.
Contacts
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SWE (RL Environments) "Reinforcement Learning"
SWE (RL Environments) "Reinforcement Learning"

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000
Remote RL Environment Engineer — Contractor (PST Overlap)
Remote RL Environment Engineer — Contractor (PST Overlap)

Open Data Science • San Francisco (CA)

Remote
USD 100,000 - 150,000
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
RLEE - Low-Level Engineering & Kernel Inference Optimization
RLEE - Low-Level Engineering & Kernel Inference Optimization

Open Data Science • San Francisco (CA)

Remote
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+2
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Software Engineer, RL Data
Software Engineer, RL Data

job-boards.greenhouse.io- JobBoard • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Full-Stack Software Engineer, Reinforcement Learning
Full-Stack Software Engineer, Reinforcement Learning

Anthropic • New York (NY)

Hybrid
USD 300,000 - 405,000
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2