Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Mercor in San Francisco is seeking an experienced Platform Engineer for our Enterprise Agent Eval Systems. You will design verifiers, build offline environments, and craft scalable grading infrastructure that works across customers, domains, and tasks.
You will collaborate with the Enterprise Platform and Applied AI teams to codify evaluation practices, reduce system friction, and enable reliable, measurable improvements in agent performance at scale.
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
Enterprise agents are complex systems, and they only pay off when their work is reliable and economically viable. Evaluation is how you get both: checking correctness is the obvious case, and routing is the subtler one, since choosing a model against cost, latency, and quality requires quality to be measurable at all. Knowing where the bar sits is the hard part. You decompose real work, take the standard from the practitioners who hold it, and encode it so an agent cannot shortcut it. You will apply what Mercor has learned building benchmarks with domain experts, and devise new methods, so that evals and rubrics keep improving and so do the agents measured against them. That work only scales with a platform behind it. This is a platform engineering role with good depth of understanding in evals. You will build the verifiers, the environments agents are measured in, and the grading infrastructure that runs at scale, abstracted across customers, domains, and tasks so that every run becomes evidence the next agent inherits instead of starting over. Read more about how we think about this: Agent Eval Systems