Turn this role into an interview — a resume and cover letter built around what this employer wants.
Acceler8 Talent is hiring a Software Engineer to design reinforcement-learning environments for frontier model labs. You will turn research objectives into scalable environments, tasks, reward signals, and evaluation rubrics that drive model learning and behavior.
You’ll work on RL environments across coding, finance, and enterprise workflows, building robust pipelines and experiments that reveal why models succeed or fail, with in-person collaboration in San Francisco.
Software Engineer, Reinforcement Learning Environments
$250k base + equity + bonus
I’m working with a highly profitable AI research infrastructure company building reinforcement-learning environments for frontier model labs.
The team creates the tasks, simulations, reward signals, evaluation rubrics, and expert trajectories used to improve how advanced models reason and act. These are not generic annotation datasets. They are structured RL environments designed to expose failure modes, measure capability, and produce useful learning signals.
This role will directly influence model post-training by turning research objectives into environments and experiments that can run at scale.
You’ll work on problems such as:
Looking for engineers who have:
This is an opportunity to build the reinforcement-learning environments and reward systems that directly shape how frontier models learn, reason, and improve.
The company operates in person from San Francisco and values speed, ownership, experimental judgment, and measurable results.