Staff ML Engineer: Agent Training & Environments

Labelbox

San Francisco (CA)

Hybrid

USD 250,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Labelbox is the RL data factory for advancing frontier agent capabilities. We build the data, environments, and evaluations that frontier labs use to train and judge their agents.

This role sits where training meets infrastructure. You will run the experiments and build the systems that run them: environments agents act in, verifiers that decide whether they succeeded, and the fine-tuning pipelines that turn that signal into a better model.

Qualifications

  • 3 year track record of shipping systems that customers and other engineers still rely on.
  • Exceptional throughput with high-quality output; your v1 is the foundation for the team.
  • Strong system and API design judgment; you make hard architectural calls and defend them.
  • Ship production code daily; you know where it breaks and how to fix it to move the team faster.
  • Build the substrate other people’s work run on—tools, CI, harnesses, libraries.

Responsibilities

  • Develop RL environments for agentic tasks, including task definitions, tool surfaces, state and reset semantics, reward design, and scalable harnesses.
  • Create verifiers or graders for open-ended work, with reliable scoring pipelines.
  • Build fine-tuning pipelines for SFT and RL, from data collection through training to evaluation.
  • Develop eval systems that measure model and product quality across agent trajectories.
  • Design and maintain training and serving infrastructure that scales with frontier labs.

Skills

Python
System design
API design
Production code
ML infrastructure
RL

Tools

GraphQL
Kubernetes
MySQL
PostgreSQL

Job description

Labelbox is the RL data factory for advancing frontier agent capabilities. We build the data, environments, and evaluations that frontier labs use to train and judge their agents.

This role sits where training meets infrastructure. You will run the experiments and build the systems that run them: environments agents act in, verifiers that decide whether they succeeded, and the fine-tuning pipelines that turn that signal into a better model.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Agent Training & Environments
Staff ML Engineer: Agent Training & Environments

EngineersOfAI • San Francisco (CA)

On-site
USD 170,000 - 260,000
Staff ML Engineer: Frontier Agent Training & Environments
Staff ML Engineer: Frontier Agent Training & Environments

B Capital • San Francisco (CA)

Hybrid
USD 250,000 - 280,000
Hybrid work model (3 days in office)
Career advancement opportunities
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

EngineersOfAI • San Francisco (CA)

On-site
USD 170,000 - 260,000
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

Labelbox • San Francisco (CA)

Hybrid
USD 250,000 - 280,000
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

B Capital • San Francisco (CA)

Hybrid
USD 250,000 - 280,000
Hybrid work model (3 days in office)
Career advancement opportunities
Staff AI Engineer - RL Environments & Agent Training
Staff AI Engineer - RL Environments & Agent Training

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 180,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+6
Forward Deployed Engineering Manager
Forward Deployed Engineering Manager

B Capital • San Francisco (CA)

On-site
USD 180,000 - 220,000
RL Environment Engineer - World-Scale Agent Training (Remote)
RL Environment Engineer - World-Scale Agent Training (Remote)

Bespoke Labs • United States

On-site
Staff Software Engineer (AI Data Platform)
Staff Software Engineer (AI Data Platform)

Labelbox • San Francisco (CA)

Hybrid
USD 250,000 - 280,000
Career advancement opportunities
Fast-paced work environment
Hybrid work model
Frontier RL Evaluation Engineer
Frontier RL Evaluation Engineer

Key Talent Solutions • San Francisco (CA)

On-site
USD 180,000 - 350,000