Member of Technical Staff, Enterprise Evals Platform

Mercor

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Relocation bonus
Housing stipend
Meals stipend
Equity grant
Free Equinox membership
Laundry reimbursement
Wellness reimbursement
Health insurance

Job summary

Mercor in San Francisco is seeking an experienced Platform Engineer for our Enterprise Agent Eval Systems. You will design verifiers, build offline environments, and craft scalable grading infrastructure that works across customers, domains, and tasks.

You will collaborate with the Enterprise Platform and Applied AI teams to codify evaluation practices, reduce system friction, and enable reliable, measurable improvements in agent performance at scale.

Qualifications

  • Experience in building evaluation systems for agent or LLM runtimes and benchmarks.
  • Ability to turn qualitative quality notions into measurable rubrics for model improvements.
  • Strong software engineering fundamentals and independent problem solving.

Responsibilities

  • Define golden sets by decomposing real tasks and encoding expert quality bars.
  • Build verifiers over agent trajectories and outputs, calibrated and hard to game.
  • Build the evaluation platform that runs offline environments and grading at scale.
  • Analyze production trajectories and turn failure modes into regression tests.
  • Run the optimization loop across models, prompts, skills, and harnesses.
  • Own rollout gates that decide whether an agent change ships.
  • Collaborate with the Enterprise Platform team and Applied AI engineers embedded with customers.

Skills

Agent engineering
Evaluation suites
LLM/agent systems
Judgment design

Tools

Harbor environments
RL environments

Job description

About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Role

Enterprise agents are complex systems, and they only pay off when their work is reliable and economically viable. Evaluation is how you get both: checking correctness is the obvious case, and routing is the subtler one, since choosing a model against cost, latency, and quality requires quality to be measurable at all. Knowing where the bar sits is the hard part. You decompose real work, take the standard from the practitioners who hold it, and encode it so an agent cannot shortcut it. You will apply what Mercor has learned building benchmarks with domain experts, and devise new methods, so that evals and rubrics keep improving and so do the agents measured against them. That work only scales with a platform behind it. This is a platform engineering role with good depth of understanding in evals. You will build the verifiers, the environments agents are measured in, and the grading infrastructure that runs at scale, abstracted across customers, domains, and tasks so that every run becomes evidence the next agent inherits instead of starting over. Read more about how we think about this: Agent Eval Systems

Responsibilities
  • Define golden sets: decompose real tasks and encode the expert quality bar.
  • Build verifiers over agent trajectories and outputs, calibrated and hard to game.
  • Build the eval platform that runs offline environments, task suites, and grading at scale.
  • Run loss analysis over production trajectories and turn failure modes into regression tests.
  • Run the optimization loop across models, prompts, skills, and harnesses.
  • Own the rollout gates that decide whether an agent change ships.
  • Partner with the Enterprise Platform team and the Applied AI engineers embedded with customers.
What We're Looking For
  • Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails.
  • Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed.
  • Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it.
  • Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems.
  • Bonus: experience with Harbor environments and RL environments.
Why Mercor
  • Impact: No agent reaches an enterprise customer without clearing the bar you set.
  • Learning: Eval work spanning frontier model measurement and live enterprise deployments, on production trajectories few teams get to see.
  • Growth: Research and systems in one role, with fast paths to owning the eval system end to end.
Benefits
  • Up to $15k relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Generous equity grant vested over 4 years
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Agents
Software Engineer, Agents

Mercor • New York (NY)

On-site
USD 150,000 - 230,000
Bi-annual performance bonus
Equity grant
Relocation bonus up to $15k
+6
Research Engineer – Benchmarking
Research Engineer – Benchmarking

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
Member of Technical Staff, Agentic Systems
Member of Technical Staff, Agentic Systems

Mercor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Generous equity
Relocation bonus
Housing bonus
+3
Member of Technical Staff, Agentic Systems
Member of Technical Staff, Agentic Systems

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Generous equity grant vested over 4 y
Relocation bonus (Bay Area)
Housing bonus (near office)
+3
Fullstack Software Engineer, Agent Platform
Fullstack Software Engineer, Agent Platform

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Bi-annual performance bonus
Equity grant
Relocation bonus up to $15k
+6
Software Engineer, Enterprise Applied AI
Software Engineer, Enterprise Applied AI

Mercor • New York (NY)

On-site
USD 140,000 - 210,000
Bi-annual performance bonus
Equity grant
Relocation bonus
+6
Fullstack Software Engineer, Agent Platform
Fullstack Software Engineer, Agent Platform

Mercor • San Francisco (CA)

On-site
USD 140,000 - 200,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
Software Engineer, Applied AI
Software Engineer, Applied AI

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
Bi-annual performance bonus
Generous equity grant
Up to relocation bonus
+6
Member of Technical Staff, Backend
Member of Technical Staff, Backend

Mercor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Generous equity grant
relocation bonus
housing bonus
+3
Software Engineer, Agents
Software Engineer, Agents

Mercor Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual performance bonus structure
Generous equity grant vested over 4 4 
Up to $15k Relocation bonus
+6