Software Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 225,000 - 275,000

Full time

34 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Acceler8 Talent is hiring a Software Engineer to design reinforcement-learning environments for frontier model labs. You will turn research objectives into scalable environments, tasks, reward signals, and evaluation rubrics that drive model learning and behavior.

You’ll work on RL environments across coding, finance, and enterprise workflows, building robust pipelines and experiments that reveal why models succeed or fail, with in-person collaboration in San Francisco.

Qualifications

  • Experience applying RL to LLM post-training, agents, or sequential decision-making.
  • Experience building environments, reward functions, evaluation tasks, or scoring systems.
  • Familiarity with RLHF, RLVR, policy optimization, or preference-based learning.
  • Ability to design controlled experiments and interpret noisy results.

Responsibilities

  • Turn research objectives into scalable RL environments and experiments.
  • Design and implement RL environments, reward signals, and evaluation rubrics.

Skills

RL in LLM post-training
Environments & rewards
RLHF/RLVR
Experiment design & analysis

Job description

Software Engineer, Reinforcement Learning Environments

$250k base + equity + bonus

I’m working with a highly profitable AI research infrastructure company building reinforcement-learning environments for frontier model labs.

The team creates the tasks, simulations, reward signals, evaluation rubrics, and expert trajectories used to improve how advanced models reason and act. These are not generic annotation datasets. They are structured RL environments designed to expose failure modes, measure capability, and produce useful learning signals.

This role will directly influence model post-training by turning research objectives into environments and experiments that can run at scale.

You’ll work on problems such as:

  • Designing RL environments for coding, finance, and enterprise workflows
  • Creating tasks that expose meaningful model and agent failure modes
  • Building reward functions and evaluation rubrics for RLHF and RLVR
  • Analyzing trajectories to understand why agents succeed or fail
  • Improving the quality and reliability of training signals
  • Measuring how environment and dataset changes affect model capability
  • Developing real-world and synthetic-data pipelines

Looking for engineers who have:

  • Experience applying RL to LLM post-training, agents, or sequential decision-making
  • Experience building environments, reward functions, evaluation tasks, or scoring systems
  • Familiarity with RLHF, RLVR, policy optimization, or preference-based learning
  • The ability to design controlled experiments and interpret noisy results

This is an opportunity to build the reinforcement-learning environments and reward systems that directly shape how frontier models learn, reason, and improve.

The company operates in person from San Francisco and values speed, ownership, experimental judgment, and measurable results.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

On-site
USD 180,000 - 220,000
RL Environments Engineer
RL Environments Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
Research Engineer
Research Engineer

Barrington James • San Francisco (CA)

On-site
USD 140,000 - 210,000
RL Environment Engineer: Shape Frontier Model Learning
RL Environment Engineer: Shape Frontier Model Learning

Acceler8 Talent • San Francisco (CA)

On-site
USD 225,000 - 275,000
Equity
Research Engineer - Post training & RL
Research Engineer - Post training & RL

techire ai • California (MO)

Hybrid
USD 180,000 - 300,000
Equity
401k
Unlimited PTO
+1
Reinforcement Learning Environment Engineer
Reinforcement Learning Environment Engineer

Open Data Science • San Francisco (CA)

Remote
USD 100,000 - 150,000
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Research Engineer, Applied AI
Research Engineer, Applied AI

HeyMilo AI • San Francisco (CA)

On-site
USD 150,000 - 210,000