Staff AI Systems Engineer: Evaluation & Post-Training Infra

NextToken AI Inc.

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

SFT & RLHF tooling
RL environments
Synthetic data pipelines
Inference optimization
Open research / OSS

Job summary

NextToken AI Inc. in San Francisco is seeking an experienced AI systems engineer to lead evaluation and post-training tooling.

You will turn production traces into golden datasets, automated graders, and regression tests, ensuring every change ships with a measurable number. You will build and maintain agentic harnesses and environments for evaluation, data generation, and RL, plus data pipelines for SFT and RL runs.

Qualifications

  • 3+ years building production AI systems. You've shipped real systems to real users.
  • Numbers over opinions. You've built eval harnesses from scratch and know the difference between a metric that hill-climbs and a metric that lies.
  • Model-agnostic. You reason about prompt vs. context vs. model changes and know when each one wins.
  • High agency. You find the highest-leverage problem and ship the fix without waiting for a spec.
  • AI-native. You use Claude, Cursor, and Codex daily and have earned opinions on where they shine and where they mislead.

Responsibilities

  • Evals as a first-class system. Turn production traces into golden datasets, automated graders, and regression tests. Every change ships with a number.
  • Agentic harnesses and environments: the tooling that runs agents reproducibly for evaluation, data generation, and RL.
  • Post-training infrastructure. Data pipelines, SFT and RL runs, and the plumbing that takes a specialized model from experiment to deployment.
  • Reward signals and verifiers. Make fuzzy notions of quality measurable and hill-climbable.
  • Latency and cost. Model routing, caching, parallel tool calls, and serving specialized models economically alongside frontier ones.

Skills

Production AI systems
Eval harnesses
Model-agnostic reasoning
High agency
Claude/Cursor/Codex usage

Job description

NextToken AI Inc. in San Francisco is seeking an experienced AI systems engineer to lead evaluation and post-training tooling.

You will turn production traces into golden datasets, automated graders, and regression tests, ensuring every change ships with a measurable number. You will build and maintain agentic harnesses and environments for evaluation, data generation, and RL, plus data pipelines for SFT and RL runs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Post-Training Evaluation & Infra
Research Engineer, Post-Training Evaluation & Infra

Ersilia • San Francisco (CA)

On-site
USD 150,000 - 350,000
Meaningful equity grants
Health, dental, and vision coverage
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Applied ML Research Intern - Post-Training & RL Experiments
Applied ML Research Intern - Post-Training & RL Experiments

NextToken AI Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 41,000 - 83,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
AI Evaluation Engineer – Reinforcement Learning & Agents
AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Enterprise AI Agent Evaluation Infrastructure Engineer
Enterprise AI Agent Evaluation Infrastructure Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 190,000
Staff AI Systems Engineer - RL Environments & Tooling
Staff AI Systems Engineer - RL Environments & Tooling

Jack & Jill • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Agent Evaluation Infrastructure Engineer
Agent Evaluation Infrastructure Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 190,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Automation Engineer: End-to-End Evaluations + Equity
AI Automation Engineer: End-to-End Evaluations + Equity

Niantic Spatial, Inc. • San Francisco (CA)

On-site
USD 158,000 - 210,000
Annual bonus
Equity
Medical coverage
+3