Research Engineer, Training and Environment Infrastructure

Ersilia

San Francisco (CA)

On-site

USD 150,000 - 350,000

Full time

33 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Meaningful equity grants
Health, dental, and vision coverage

Job summary

Ersilia in San Francisco is building the data and infrastructure behind frontier model training and deployed AI agents. You will own evaluation systems, ranking training tasks by value and catching drift as data and tasks evolve.

This role is hands-on and in person, six days a week, requiring strong systems engineering, RL experience, and the ability to ship production-grade evaluators tied to training pipelines.

Qualifications

  • Experience building environments, evaluations, or RL pipelines in production.
  • Hands-on RL research or data science experience.
  • Strong systems engineering skills and the ability to ship end-to-end evaluators tied to training pipelines.

Responsibilities

  • Rank training value before spending compute and build a system to rank tasks by expected training value.
  • Design scoring that runs cheaply at scale, tuned to customer-specific outcomes.
  • Detect and adapt to distribution drift by updating evaluations as customer requests evolve.

Skills

RL pipelines
Systems engineering
Reward design
Ambiguous evaluation problems

Job description

About the company

We work on the data and infrastructure behind frontier model training and deployed AI agents. We partner with AI labs and large organizations on post-training research, infrastructure, and deployments. We are an early-stage, revenue-generating team hiring for in-person roles in San Francisco.

About the role

You sit at the center of the post-training loop. Before compute is spent on a dataset, you decide whether it is worth it. After a model change ships, you determine whether it helped.

The work spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic and task shapes evolve. This is not a fixed benchmark you maintain once and forget.

What you'll do
Rank training value before spending compute

Work out which tasks in a dataset are worth training on, and build a system that ranks every task by expected training value.

Build scoring that runs cheaply at scale

Design automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good outcome rather than generic correctness.

Catch drift before it becomes a problem

Notice when real customer requests have moved far enough from the existing test set that it no longer describes the job, then rebuild the evaluations to match.

What we're looking for
  • You have built environments, evaluations, or RL pipelines that ran in production.
  • You have built RL data or done hands-on RL research, rather than working adjacent to it.
  • You are a strong systems engineer who can pick up any stack.
  • You have substantive opinions on reward design, including what makes a checker trustworthy and how models learn to game them.
  • You are comfortable owning ambiguous evaluation problems end to end, from building the first version and inspecting the data to iterating until the system is useful.
  • You can work six days a week, in person, in San Francisco.
Nice to have
  • Experience with LLM-as-judge systems, automated graders, or rubric design for model outputs.
  • A background in applied ML research, research engineering, or data science at a lab or fast-moving startup.
  • Experience building evaluation infrastructure tied directly to a training pipeline, not just offline benchmarking.
Compensation and benefits

$150,000 to $350,000 USD base, depending on experience and seniority

  • Meaningful equity grants for early team members
  • Health, dental, and vision coverage
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Scientist/Engineer
Research Scientist/Engineer

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Staff MLE, Reinforcement Learning
Staff MLE, Reinforcement Learning

People In AI • San Francisco (CA)

On-site
USD 270,000 - 280,000
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2