AfterQuery — Software Engineer, RL Environments

davidjoseph-co

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive equity

Job summary

AfterQuery in San Francisco, CA is seeking a Software Engineer for RL Environments to design data slices and evaluation rubrics powering frontier model training.

You will own end-to-end data pipelines, experiment with RLHF/RLVR signals, and collaborate with research teams to translate training objectives into concrete data and evaluation specs. 1–4 years of software engineering experience required; on-site work. Competitive compensation with equity and rapid impact in a small, early-stage team.

Qualifications

  • 1–4 years of software engineering experience.
  • Design targeted data slices to surface model failure modes across domains.
  • Build and iterate evaluation rubrics powering RLHF/RLVR pipelines.
  • Develop quantitative frameworks to measure dataset quality, diversity, and impact on model alignment.
  • Own end-to-end real-world and synthetic data pipelines, from scoping to production-ready evaluation specs.
  • Run annotator modeling experiments to improve model capabilities.

Responsibilities

  • Design data slices and explore data shapes that expose meaningful model failure modes across domains.
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines.
  • Model annotator behavior and run experiments to improve model capabilities.
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability.
  • Create and manage both real-world and synthetic data pipelines.
  • Partner with lab research teams to translate training objectives into concrete data and evaluation specs.

Skills

Software engineering
Data pipelines
RLHF/RLVR
Experimentation
Model evaluation

Job description

AfterQuery — Software Engineer, RL Environments

Type: Full-time | On-site | San Francisco, CA Compensation: $180,000–$220,000 + competitive equity (see comp note below) Hiring count: 3 Visa sponsorship: O-1, OPT Reports to: Not specified (initial screen: Alec; second round: Michael/Sam)

About AfterQuery

AfterQuery builds the training data and evaluation infrastructure that frontier AI labs use to improve their models, designing high‑signal datasets and running rigorous evaluations that go beyond static benchmarks. It's a small, early team (post–Series A) where individual contributors have direct impact on how the next generation of models learns. The founding team comes from Jane Street, Citadel, Google, Goldman, and Stanford AI Lab.

Founded: 2025 | Team size: 11–50 | Total funding: $30M raised (~$300M valuation) | Industry: Consumer Tech | Website: afterquery.com | Office: San Francisco, CA

Why Candidates Should Join
  • Outsized total cash: $200K base plus profit share of roughly 150% of base, bringing expected total cash to around $500K, plus competitive equity.
  • Direct line to frontier models: Output feeds directly into model training runs at scale, working hands‑on with research teams at top AI labs.
  • High ownership, early team: Small post–Series A team where ICs scope, build, and ship end to end.
The Role

As a SWE (Environments), you design the datasets and evaluation rubrics that directly influence how frontier models learn — going from hypothesis to live experiment quickly, with output feeding directly into model training.

What You'll Be Doing
  • Design data slices and explore data shapes that expose meaningful model failure modes across domains like finance, code, and enterprise workflows.
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines.
  • Model annotator behavior and run experiments to improve different model capabilities.
  • Develop quantitative frameworks for measuring dataset quality, diversity, and downstream impact on model alignment and capability.
  • Create and manage both real‑world and synthetic data pipelines.
  • Partner with lab research teams to translate their training objectives into concrete data and evaluation specifications.

Tech stack: Not specified

Requirements
  • 1–4 years of software engineering experience with strong technical depth.
  • Design targeted data slices that surface model failure modes across high‑stakes domains (finance, code generation, enterprise workflows).
  • Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR training pipelines.
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability.
  • Own end‑to‑end real‑world and synthetic data pipelines, from scoping with research teams to production‑ready evaluation specs.
  • Run annotator modeling experiments to improve model capabilities across task types.
Green Flags
  • Experience at RL environment companies.
  • Background in AI safety or benchmarking organizations like METR or Artificial Analysis.
  • Genuine obsession with how data structure, selection, and quality drive model behavior.
  • Ability to design lightweight experiments and move fast.
  • Former founders or early engineers at early stage startups.
  • Demonstrated ability to work hard, learn fast, and care deeply about details.
Red Flags
  • Pure research profile with limited engineering output; this is a SWE role, shipping matters.
  • Looking for standard product engineering work — the real scope is data pipelines, reward modeling, and eval infra.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - RL Environments San Francisco
Software Engineer - RL Environments San Francisco

AfterQuery • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 210,000
Equity
Growth opportunities
Software Engineer, RL Environments
Software Engineer, RL Environments

David Joseph & Company • San Francisco (CA)

On-site
USD 180,000 - 220,000
AfterQuery — Research Scientist, Post-Training
AfterQuery — Research Scientist, Post-Training

davidjoseph-co • San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity
On-site work
Software Engineer, Platform / Applied AI (Full-Stack)
Software Engineer, Platform / Applied AI (Full-Stack)

davidjoseph-co • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive salary + equity (4-year)
Daily lunch & dinner
Gym membership
+1
Research Scientist - Frontier Data
Research Scientist - Frontier Data

AfterQuery • San Francisco (CA)

On-site
USD 250,000 - 450,000
AfterQuery — Strategic Projects Associate
AfterQuery — Strategic Projects Associate

davidjoseph-co • San Francisco (CA)

On-site
USD 80,000 - 110,000
Equity
Competitive base salary
On-site in San Francisco
AfterQuery — Strategic Projects Lead
AfterQuery — Strategic Projects Lead

davidjoseph-co • San Francisco (CA)

On-site
USD 130,000 - 150,000
Equity
Profit-sharing on contracts
AfterQuery — Strategic Projects Lead - Coding
AfterQuery — Strategic Projects Lead - Coding

davidjoseph-co • San Francisco (CA)

On-site
USD 130,000 - 150,000
Equity
On-site work in San Francisco
Strategic Projects Lead San Francisco
Strategic Projects Lead San Francisco

AfterQuery • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Technical Strategic Projects Lead
Technical Strategic Projects Lead

AfterQuery • San Francisco (CA)

On-site
USD 300,000
Equity opportunities
Competitive salary
Work with a renowned team