RL Environments Engineer

Bespoke Labs Inc.

Mountain View, Northern (CA, KY)

Hybrid

USD 250,000 - 300,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
Visa sponsorship and relocation relief
Direct impact on industry training and

Job summary

Bespoke Labs Inc. seeks an engineer to build the machinery that turns environment ideas into hundreds or thousands of validated agentic coding tasks.

You will design pipelines that mass-produce environments, craft complex coding worlds around real codebases, and push throughput with minimal manual work per task. We measure you on volume and quality of environments shipped, not papers, and you will directly influence frontier coding agents and their validation, with strong emphasis on scalable

Qualifications

  • Experience building agentic coding tasks or environments and quantifying volume and cost.
  • Experience scaling output through automation rather than more people.
  • Strong software engineering fundamentals and familiarity with production-grade code in multiple languages.

Responsibilities

  • Build environment-generation pipelines to produce RL environments at scale with templating, grading, verification, and QA.
  • Create high-fidelity coding worlds around real codebases with proper tooling and dependencies.
  • Scale task creation to thousands with automation handling the heavy lifting.
  • Build internal tools and infrastructure to raise throughput and remove bottlenecks.
  • Own the full task lifecycle from prompt to evaluation, ensuring rigor and fairness.
  • Defend quality at scale by catching reward hacking and grader loopholes.

Job description

About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

About the Role
This is a delivery role. We want an engineer who has built the machinery that turns environment ideas into hundreds or thousands of validated agentic coding tasks, and who can do it here, fast.
You will not be studying environments in the abstract. You will build the pipelines that mass-produce them, design the complex coding worlds agents train inside, and keep pushing throughput: more environments, higher quality, less manual work per task. We will measure you on the volume and quality of environments you ship, not on papers.
The thing we care about most is whether you have done this before. If you have stood up an environment-generation pipeline, scaled agentic task creation into the hundreds or thousands, and shipped it, we want to talk.

What You'll Do
  • Build environment-generation pipelines. Own the systems that produce RL environments programmatically, including templating, automated grading, verification, and QA, so the team ships environments at scale instead of one at a time.

  • Create complex coding worlds. Build high-fidelity environments around real codebases, with the conventions, dependencies, tooling, and technical debt that real software actually has.

  • Scale agentic task creation to thousands. Take task generation from handfuls to hundreds and thousands of validated agentic coding tasks, with automation doing the heavy lifting.

  • Build tools that raise throughput. Find the bottlenecks in environment production and remove them. Build the internal tooling and infrastructure that makes everyone on the team faster.

  • Own the full task lifecycle. Prompt, environment, grader, running frontier models against the task, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.

  • Defend quality at scale. Catch reward hacking and grader loopholes, and build the verification and standards that hold the bar as volume grows.

  • Direct coding agents heavily. Use frontier coding agents to build and validate environments faster, judging their output and catching the subtle failures.

Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.

What We're Looking For

A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.

Experience scaling that output through automation rather than through more people doing more manual work.

Strong software engineering fundamentals and fluency in several languages that holds up in production code.

Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.

An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.

A clear sense of what frontier coding agents can and cannot do, and where they cut corners.

Ownership. You build, debug, and ship without much supervision.

You May Be a Good Fit If You Also
  • Have worked on RL training systems, post-training, verifiers, or tool-use harnesses

  • Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure

  • Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment

  • Have contributed to a public agentic benchmark such as Terminal-Bench

  • Have open-source work that other people depend on

What We Offer
  • Location: Mountain View, CA (Onsite)

  • Base Salary: $250,000 – $300,000 USD / year

  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks

  • Health, dental, and vision coverage

  • 401(k)

  • Daily onsite lunch provided

  • Visa sponsorship and relocation support available

  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
RL Environments Engineer
RL Environments Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
Research Engineer
Research Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health coverage
Opportunity to work with leading AI labs
Competitive salary and equity
Backend Engineer
Backend Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 100,000 - 140,000
Health coverage
Opportunity to work with leading AI research labs
Backend Engineer
Backend Engineer

bespokelabs • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Competitive salary and equity
Opportunity to work with leading AI research labs
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
Member of Technical Staff, RL Systems
Member of Technical Staff, RL Systems

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Location-based hybrid policy
Research Engineer, Frontier Evals & Environments
Research Engineer, Frontier Evals & Environments

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
Staff Software Engineer, RL Environments
Staff Software Engineer, RL Environments

Scale AI • San Francisco (CA), New York (NY)

On-site
USD 252,000 - 315,000
Health, dental, vision
Equity
Generous PTO
+1
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

Hybrid
USD 180,000 - 220,000