RL Data Engineer - Design Training Tasks for Coding Agents

Anysphere

San Francisco, Northern (CA, KY)

Hybrid

USD 130,000 - 180,000

Full time

31 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anysphere is seeking a Software Engineer, Reinforcement Learning to join our RL Data team and help design tasks, rewards, and environments that train coding agents. You will influence how we measure progress, curate data, and push the boundaries of what our models can learn.

You will iterate on traces and failure modes, turning insights into reusable pipelines and environments for multiple teams across the company.

Qualifications

  • You write careful, fast code and have strong software engineering fundamentals.
  • You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
  • You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
  • You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.

Responsibilities

  • Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
  • Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
  • Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
  • Partnering with research on whether a dataset is actually teaching the thing we think it is.

Job description

Anysphere is seeking a Software Engineer, Reinforcement Learning to join our RL Data team and help design tasks, rewards, and environments that train coding agents. You will influence how we measure progress, curate data, and push the boundaries of what our models can learn.

You will iterate on traces and failure modes, turning insights into reusable pipelines and environments for multiple teams across the company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer: RL Environments & AI Task Design
Senior Software Engineer: RL Environments & AI Task Design

Mechanize, Inc. • San Francisco (CA)

On-site
USD 360,000 - 440,000
Health insurance
Dental insurance
Vision insurance
+3
RL Research Scientist - Coding Agents & Data-Driven Impact
RL Research Scientist - Coding Agents & Data-Driven Impact

Cursor • California (MO)

On-site
USD 140,000 - 230,000
Remote AI Training Engineer for RL Software Tasks
Remote AI Training Engineer for RL Software Tasks

YO AI Labs • Maryland

Remote
USD 69,000 - 165,000
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff Software Engineer — RL Infrastructure & Platforms
Staff Software Engineer — RL Infrastructure & Platforms

Anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior AI RL Engineer — Systems & Infrastructure
Senior AI RL Engineer — Systems & Infrastructure

Mechanize • Oakland (CA), San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1
RL Environment Engineer - World-Scale Agent Training (Remote)
RL Environment Engineer - World-Scale Agent Training (Remote)

Bespoke Labs • United States

On-site
Remote Senior Software Engineer — RL Environment Designer
Remote Senior Software Engineer — RL Environment Designer

YO AI Labs • Los Angeles (CA)

Remote
USD 40,000 - 70,000
Junior Reinforcement Learning Engineer
Junior Reinforcement Learning Engineer

Mechanize • Oakland (CA), San Francisco (CA)

On-site
USD 90,000 - 150,000
Health insurance
Dental insurance
Vision insurance
+1
Staff Research Engineer: Coding AI Environments
Staff Research Engineer: Coding AI Environments

Cerebras • California (MO)

On-site
USD 250,000 - 350,000