Staff Engineer, Enterprise Evaluation Platform

Doist

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation bonus
Housing stipend
Meals stipend
Equity grant
Free Equinox membership
Laundry reimbursement
Wellness reimbursement
Health insurance

Job summary

Mercor in San Francisco is seeking an experienced Platform Engineer for our Enterprise Agent Eval Systems. You will design verifiers, build offline environments, and craft scalable grading infrastructure that works across customers, domains, and tasks.

You will collaborate with the Enterprise Platform and Applied AI teams to codify evaluation practices, reduce system friction, and enable reliable, measurable improvements in agent performance at scale.

Qualifications

  • Experience in building evaluation systems for agent or LLM runtimes and benchmarks.
  • Ability to turn qualitative quality notions into measurable rubrics for model improvements.
  • Strong software engineering fundamentals and independent problem solving.

Responsibilities

  • Define golden sets by decomposing real tasks and encoding expert quality bars.
  • Build verifiers over agent trajectories and outputs, calibrated and hard to game.
  • Build the evaluation platform that runs offline environments and grading at scale.
  • Analyze production trajectories and turn failure modes into regression tests.
  • Run the optimization loop across models, prompts, skills, and harnesses.
  • Own rollout gates that decide whether an agent change ships.
  • Collaborate with the Enterprise Platform team and Applied AI engineers embedded with customers.

Skills

Agent engineering
Evaluation suites
LLM/agent systems
Judgment design

Tools

Harbor environments
RL environments

Job description

Mercor in San Francisco is seeking an experienced Platform Engineer for our Enterprise Agent Eval Systems. You will design verifiers, build offline environments, and craft scalable grading infrastructure that works across customers, domains, and tasks.

You will collaborate with the Enterprise Platform and Applied AI teams to codify evaluation practices, reduce system friction, and enable reliable, measurable improvements in agent performance at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, Enterprise AI Infrastructure
Senior Platform Engineer, Enterprise AI Infrastructure

Doist • New York (NY), San Francisco (CA)

On-site
USD 170,000 - 250,000
Up to $15k relocation bonus
$10K housing bonus
$1.5K monthly meals stipend
+5
Fullstack Engineer - Enterprise AI Agent Platform (SF)
Fullstack Engineer - Enterprise AI Agent Platform (SF)

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Bi-annual performance bonus
Equity grant
Relocation bonus
+6
Member of Technical Staff, Enterprise Evals Platform
Member of Technical Staff, Enterprise Evals Platform

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation bonus
Housing stipend
Meals stipend
+5
Senior Platform Architect – Enterprise AI Infrastructure
Senior Platform Architect – Enterprise AI Infrastructure

Apply • New York (NY)

On-site
USD 170,000 - 260,000
Relocation bonus
Housing bonus
Meal stipend
+7
Senior Software Engineer, Agent Evaluation Platform
Senior Software Engineer, Agent Evaluation Platform

Moveworks • Mountain View (CA)

Hybrid
USD 161,000 - 274,000
Full-Stack Engineer — AI Assessment Platform, Equity
Full-Stack Engineer — AI Assessment Platform, Equity

Mercor • San Francisco (CA)

On-site
USD 140,000 - 230,000
Bi-annual performance bonus
Equity grant (4 years)
Relocation bonus up to $15k
+6
Full-Stack Engineer, AI Agent Platform
Full-Stack Engineer, AI Agent Platform

Mercor • San Francisco (CA)

On-site
USD 140,000 - 200,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
Open Source Evaluation Engineer
Open Source Evaluation Engineer

Mercor • United States

Remote
USD 90,000 - 150,000
AI Benchmarking & Evaluation Engineer
AI Benchmarking & Evaluation Engineer

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
Fullstack Engineer, Agent Platform - Build AI Workflows
Fullstack Engineer, Agent Platform - Build AI Workflows

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Bi-annual performance bonus
Equity grant
Relocation bonus up to $15k
+6