Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Ersilia in San Francisco is building the data and infrastructure behind frontier model training and deployed AI agents. You will own evaluation systems, ranking training tasks by value and catching drift as data and tasks evolve.
This role is hands-on and in person, six days a week, requiring strong systems engineering, RL experience, and the ability to ship production-grade evaluators tied to training pipelines.
We work on the data and infrastructure behind frontier model training and deployed AI agents. We partner with AI labs and large organizations on post-training research, infrastructure, and deployments. We are an early-stage, revenue-generating team hiring for in-person roles in San Francisco.
You sit at the center of the post-training loop. Before compute is spent on a dataset, you decide whether it is worth it. After a model change ships, you determine whether it helped.
The work spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic and task shapes evolve. This is not a fixed benchmark you maintain once and forget.
Work out which tasks in a dataset are worth training on, and build a system that ranks every task by expected training value.
Build scoring that runs cheaply at scaleDesign automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good outcome rather than generic correctness.
Catch drift before it becomes a problemNotice when real customer requests have moved far enough from the existing test set that it no longer describes the job, then rebuild the evaluations to match.
$150,000 to $350,000 USD base, depending on experience and seniority