Research Engineer - Evals

AGI, Inc.

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive cash and equity
Top-tier relocation support
In-person work environment

Job summary

AGI, Inc. in San Francisco is looking for a skilled professional to build evaluation harnesses that ensure models and agents are performing at their best. You will set criteria for releases, audit existing evaluation processes, and develop tooling to assist research and product teams.

The position emphasizes collaboration and delivery of dependable performance metrics to improve AI capabilities. You'll need to have a firm grasp on non-deterministic systems and QA processes in a production environment, enhancing user experience and device performance.

Competitive compensation and relocation support are offered.

Qualifications

  • Ability to audit evals and identify gaps.
  • Experience with agent performance evaluation.
  • Skilled in measuring long-horizon tasks.

Responsibilities

  • Build eval harness for model capability and user experience.
  • Create dashboards for researcher experiment loops.
  • Define readiness criteria for product launches.

Skills

Agent evaluation
Tool use measurement
QA at OEM scale
Multilingual behavior assessment

Job description

Think Different. Build the Future.
Our Mission

Build everyday AGI. Trustworthy, consumer-grade agents that redefine human–AI collaboration for millions. Software shouldn’t wait for commands; it should partner with you, amplifying what you can do every single day.

Why AGI, Inc.

We’re a stealth team of elite founders and AI researchers, with backgrounds spanning Stanford, OpenAI, and DeepMind. We’re industry leaders in mobile and computer-use agents, bringing these capabilities to consumer scale.

Grounded in years of agent research, our AI is designed with trustworthiness and reliability as core pillars, not afterthoughts.

We are supported by tier-1 investors who funded the first generation of AI giants; now they’re backing us to build the next: everyday AGI. (Watch the demo)

If you see possibility where others see limits, read on.

You decide what "better" means.

Models, agents, and product features all ship behind one question: did this actually get better? Without a strong evals function, the lab ships vibes. With one, every training run, every prompt change, every agent capability moves a number we trust — and the team makes decisions on real signal, not the loudest opinion in the room.

You’ll build the eval harness for AGI — across model capability, agentic behavior, on-device performance, and end-user experience. You’ll set the bar for what counts as "shipped" and protect it from the gravity of product deadlines.

Tasks you will own
  • The eval suites that gate every model and agent release — capability, behavior, regressions, and human-rated rubrics that catch what automated evals miss
  • The dashboards and tooling that make researcher experiment loops fast and leadership decisions easy
  • The bar — what counts as ready to ship, and how we know
Areas where you will assist
  • Research, by making sure what we measure is what we want
  • Product engineers, by instrumenting real-user behavior on real devices
  • Partnerships, by translating "did it get better" into language an OEM partner can hold us to
Skills you’ll be expected to teach
  • How to measure non-deterministic systems — agent eval, tool use, long-horizon tasks, multilingual behavior
  • How to push back on a metric that’s being gamed without breaking the team
Skills you’ll be expected to learn
  • On-device perf trade-offs and how they show up in real-user evals
  • What QA-ing AI at OEM scale actually looks like
  • The realities of shipping consumer agents to production partners
Timeline of success

After 30 days — You’ve audited every eval we run today and produced a sharp doc on what’s good, what’s noise, and what’s missing. You’ve fixed the most embarrassing gap.

After 60 days — You’ve stood up a new eval surface — agentic, on-device, or behavioral — and the team is making real decisions on its output. Researchers come to you before launching a run, not after.

After 90 days — Releases now ship against your eval bar, not a vibe-check. You’ve caught a regression that would have shipped, and cleared a launch the team was nervous about. You’re shaping the research roadmap by surfacing where we’re flat, where we’re climbing, and where we’re lying to ourselves.

Compensation

Competitive cash and meaningful equity. Top-tier relocation and immigration support. SF, in person.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Evals
Research Engineer - Evals

Pantera Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash compensation
Equity opportunities
Relocation support
+1
Research Engineer - Evals
Research Engineer - Evals

agi-inc • Santa Fe (NM)

On-site
USD 140,000 - 200,000
Relocation support
Immigration support
In-person SF
Research Engineer - Evals
Research Engineer - Evals

AGI Inc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash and equity
Top-tier relocation support
AI Researcher
AI Researcher

agi-inc • Santa Fe (NM)

On-site
USD 180,000 - 240,000
Relocation support
Meaningful equity
In-person SF relocation and visa
AI Engineer - Backend
AI Engineer - Backend

AGI, Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and meaningful equity
Top-tier relocation and immigration support
Agent Post-Training, Frontier Evals and Environments Research
Agent Post-Training, Frontier Evals and Environments Research

United States Digital Space LLC • San Francisco (CA)

On-site
USD 120,000 - 160,000
Evals Infrastructure Tech Lead / Manager
Evals Infrastructure Tech Lead / Manager

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
AI Researcher / Engineer / Intern
AI Researcher / Engineer / Intern

Egra • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Platinum-tier health insurance
Uncapped compute access
ML Platform & Infrastructure Engineer
ML Platform & Infrastructure Engineer

agi-inc • Santa Fe (NM)

On-site
USD 110,000 - 170,000
Medical insurance
Dental insurance
Vision insurance
+2
Research Engineer, Frontier Evals & Environments
Research Engineer, Frontier Evals & Environments

OpenAI • Los Angeles (CA)

On-site
USD 100,000 - 150,000