Research Engineer – Evals

Rnb Consultancy

San Francisco (CA)

On-site

USD 200,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Relocation support
Visa/immigration assistance

Job summary

San Francisco AI Lab is hiring to build and own evaluation ecosystems for on-device agents and model releases. You will define readiness criteria, design human-scored rubrics, and ship dashboards and tooling to speed research and leadership decisions.

Ideal candidates excel at engineering instrumentation, rapid experiment cycles, and communicating what constitutes meaningful improvement under deadline pressure.

Qualifications

  • Design and maintain evaluation harnesses for AI models and agents.
  • Create robust rubrics and ensure evaluation quality under tight deadlines.
  • Build dashboards and tooling to accelerate research iterations and leadership decisions.

Responsibilities

  • Own the eval suites for model and agent releases (capabilities, behavior, regressions).
  • Define what "ready to ship" means with concrete evidence.
  • Bridge research, product engineering, and partnerships to translate improvements into verifiable commitments.

Skills

Eval harnesses
Dashboards
Instrumentation
Rubrics design
On-device evaluation

Tools

Dashboards tooling
Experiment tooling

Job description

\"Did it actually get better?\" If you want to be the one who answers that with numbers the whole team trusts, read on.


A San Francisco AI lab is building consumer-grade agents that run privately on phones, laptops and other devices. They've held the top spot on the industry's leading mobile-agent benchmark since late 2025, work with a leading mobile chipmaker and a global device maker, and are backed by top-tier VCs.


Without a serious evals function, releases are based on gut feeling. With one, the team knows whether each training run or prompt tweak helped, and trusts the number. You build that function: evals covering what the models can do, how the agents behave, how they perform on device, and what users actually experience. And you decide what \"shipped\" means, then hold that line when deadlines push.


What you'll own


  • The eval suites every model and agent release has to pass: capabilities, behavior, regressions, plus rubrics scored by people for everything automated checks can't see

  • Dashboards and tools that speed up research iterations and make leadership decisions simple

  • The definition of \"ready to ship\", and the evidence behind it

  • The bridge to research (measuring what they really care about), product engineers (tracking how real users behave on real hardware) and the partnerships team (translating \"it improved\" into commitments a hardware partner can verify)


Your first 90 days


  • Day 30: you've reviewed every existing eval, written up what's useful, what's noise and what's missing, and closed the biggest gap

  • Day 60: you've launched a new area of evaluation

  • Day 90: releases go out against your standard, and you've stopped at least one regression from shipping


What you bring


  • Eval harnesses for systems that never give the same answer twice: agents, tool calling, multi-step and long-horizon tasks, behavior across languages

  • Strong engineering and tooling skills: dashboards, instrumentation, quick experiment cycles

  • You design human-scored rubrics, and when a metric gets gamed you call it out without wrecking team morale

  • Real urgency. You hold the standard under deadline pressure and say \"not ready\" when it isn't.

  • Evals built in a real setting: a data or eval company, or a frontier lab's eval team

  • Hypergrowth, founder or early-hire experience is a plus; relevant work beats a famous name


Bonus points


  • Agentic, tool-use or on-device evals

  • Consumer AI or QA at device-maker scale


What's in it for you


  • $200k–$250k base + equity

  • You define what \"better\" means at a frontier on-device lab

  • Relocation and immigration support


Good to know


  • Full-time, in person in San Francisco, 9-9-6

  • Process: intro → quick talk with the founder → technical conversation (no live coding, no puzzles) → half-day onsite → work trial → offer

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer - Evals
Research Engineer - Evals

AGI, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash and equity
Top-tier relocation support
In-person work environment
Research Engineer - Evals
Research Engineer - Evals

AGI Inc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash and equity
Top-tier relocation support
AI Researcher – RL & On-Device Models
AI Researcher – RL & On-Device Models

Rnb Consultancy • San Francisco (CA)

On-site
USD 200,000 - 250,000
Relocation support
Immigration support
Evaluation Engineer (Consulting / Contract-to-Hire) Seattle/Remote
Evaluation Engineer (Consulting / Contract-to-Hire) Seattle/Remote

AI Ethics Network • Northern (KY)

Hybrid
USD 138,000 - 220,000
Founding AI Engineer, Agents & Evaluation
Founding AI Engineer, Agents & Evaluation

Zingage • New York (NY)

On-site
USD 190,000 - 240,000
Competitive base and meaningful equity
Equipment stipend
Luxury gym membership in NYC
+4
Member of Technical Staff, Evals Lead
Member of Technical Staff, Evals Lead

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical package
Dental package
+8
Head of Evals, AI Red Teaming
Head of Evals, AI Red Teaming

Trajectory Labs, PBC • Berkeley (CA), Northern (KY)

On-site
USD 200,000 - 400,000
Health coverage stipend
401(k)
Visa sponsorship
+1
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
Software Engineer - Evals
Software Engineer - Evals

Maven Ventures • Palo Alto (CA)

On-site
USD 175,000 - 275,000
Equity
Medical coverage
401(k) retirement plan
+3
AI Researcher — Browser Agents & Product Impact
AI Researcher — Browser Agents & Product Impact

MULTI·ON • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive compensation package
Significant equity
Access to substantial compute resources