Software Engineer, RL Environments

David Joseph & Company

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

David Joseph & Company in San Francisco, CA is seeking a SWE (Environments) to design datasets and evaluation rubrics that influence frontier model training. You will move quickly from hypotheses to live experiments, with outputs feeding model training at scale.

The role demands 1–4 years of software engineering experience, ownership of end‑to‑end data pipelines, and the ability to design lightweight experiments across finance, code, and enterprise workflows.

Qualifications

  • 1–4 years of software engineering experience with strong technical depth.
  • Design targeted data slices that surface model failure modes across high‑stakes domains.
  • Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR pipelines.
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability.
  • Own end‑to‑end real world and synthetic data pipelines from scoping with research teams to production‑ready evaluation specs.

Responsibilities

  • Design data slices and explore data shapes to reveal model failure modes in domains like finance, code, and enterprise workflows.
  • Build and refine evaluation rubrics and reward signals for RLHF/RLVR training pipelines.
  • Model annotator behavior and run experiments to improve model capabilities.
  • Develop quantitative frameworks to measure dataset quality and downstream impact on alignment and capability.
  • Create and manage both real‑world and synthetic data pipelines.

Skills

Software engineering
Data pipelines
Experiment design
RLHF pipelines
Model evaluation

Job description

San Francisco, CA · On-site · Full-time
Compensation: $180,000–$220,000 + competitive equity

About the Company

An early-stage (post–Series A) company building the training data and evaluation infrastructure that frontier AI labs use to improve their models — designing high-signal datasets and running rigorous evaluations that go beyond static benchmarks. A small team where individual contributors have direct impact on how the next generation of models learns. The company has raised $30M (~$300M valuation), with a founding team drawn from Jane Street, Citadel, Google, Goldman, and Stanford AI Lab.

Founded 2025 · 11–50 people · Industry: Consumer Tech

The Role

As a SWE (Environments), you'll design the datasets and evaluation rubrics that directly influence how frontier models learn — going from hypothesis to live experiment quickly, with output feeding directly into model training runs at scale.

What you'll be doing

  • Design data slices and explore data shapes that expose meaningful model failure modes across domains like finance, code, and enterprise workflows
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines
  • Model annotator behavior and run experiments to improve different model capabilities
  • Develop quantitative frameworks for measuring dataset quality, diversity, and downstream impact on model alignment and capability
  • Create and manage both real‑world and synthetic data pipelines
  • Partner with lab research teams to translate their training objectives into concrete data and evaluation specifications

Tech stack: Not specified

Requirements
  • 1–4 years of software engineering experience with strong technical depth
  • Design targeted data slices that surface model failure modes across high‑stakes domains (finance, code generation, enterprise workflows)
  • Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR training pipelines
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability
  • Own end‑to‑end real world and synthetic data pipelines, from scoping with research teams to production‑ready evaluation specs
  • Run annotator modeling experiments to improve model capabilities across task types
Green Flags
  • Experience at RL environment companies
  • Background in AI safety or benchmarking organizations like METR or Artificial Analysis
  • Genuine obsession with how data structure, selection, and quality drive model behavior
  • Ability to design lightweight experiments and move fastFormer founders or early engineers at early stage startups
  • Demonstrated ability to work hard, learn fast, and care deeply about details
Red Flags
  • Pure research profile with limited engineering output, this is a SWE role, shipping matters
  • Looking for standard product engineering work — the real scope is data pipelines, reward modeling, and eval infra
Why Join
  • Outsized total cash: base plus substantial profit share, plus competitive equity
  • Direct impact on frontier AI model development, working with the world's leading AI labs
  • High ownership on a small, early team — scope, build, and ship end to end
Details
  • Location: San Francisco, CA
  • Work policy: On-site
  • Compensation: $180,000–$220,000 + equity
  • Visa sponsorship: O-1, OPT
  • Employment type: Full-time
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AfterQuery — Software Engineer, RL Environments
AfterQuery — Software Engineer, RL Environments

davidjoseph-co • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive equity
Software Engineer - RL Environments San Francisco
Software Engineer - RL Environments San Francisco

AfterQuery • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 210,000
Equity
Growth opportunities
SWE (RL Environments) "Reinforcement Learning"
SWE (RL Environments) "Reinforcement Learning"

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000
Member of Technical Staff, Platform Engineering
Member of Technical Staff, Platform Engineering

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 250,000
Healthcare
Relocation support
401k with 4% match
+3
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
Software Engineer, Platform / Applied AI
Software Engineer, Platform / Applied AI

David Joseph & Company • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive salary + equity
Daily lunch & dinner
Free gym membership
+1
Research Scientist, Post-Training
Research Scientist, Post-Training

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 450,000
Software Engineer, RL Environments & Evaluation Pipelines
Software Engineer, RL Environments & Evaluation Pipelines

davidjoseph-co • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive equity
Junior Software Engineer
Junior Software Engineer

Simplify • San Francisco (CA)

On-site
USD 255,000 - 345,000
Senior AI Forward Deployed Engineer
Senior AI Forward Deployed Engineer

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2