Evals Engineer

Zof AI

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

MacBook Pro
Claude Code Ultra

Job summary

Zof AI in San Francisco, CA is seeking an Evals Engineer to design and build the tests that determine whether AI products actually work. You’ll create eval suites, verification harnesses, and quality gates that confirm the system ships correct behavior, combining QA discipline with domain judgment.

You’ll join a fast-moving, high-performance team, work closely with engineering and product, and drive measurable improvements in AI reliability, safety, and user impact through rigorous evaluation

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

QA testing
LLM evaluation
Automation scripting
Analytical thinking
Communication skills
Ownership

Tools

CI/CD tools
Test automation

Job description

Compensation

Competitive salary

  • Claude Code Ultra

San Francisco, CA

Full-time

Mid to Senior

On-site

Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for candidates who want to contribute directly to the development of ambitious AI products in a high-performance environment.

Must understand how to measure whether AI systems actually work, beyond demos.

About This Role

Zof AI is seeking an Evals Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Responsibilities
  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.
Requirements
  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.
Nice to have
  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.
What we provide in San Francisco
  • MacBook Pro
  • Premium AI development tools
  • Cursor Ultra
  • Claude Code Ultra
  • OpenAI Codex Max or equivalent advanced AI tooling
  • Access to a high-performance AI product environment
  • Close collaboration with leadership, engineering, and customers
  • Opportunity to work in the San Francisco AI ecosystem
  • Wellness and productivity support where applicable
  • Competitive startup environment
  • High ownership
  • Direct product impact

Benefits may depend on role and final offer terms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer
Machine Learning Engineer

Zof AI • San Francisco (CA)

On-site
USD 140,000 - 210,000
MacBook Pro
Premium AI tools
Cursor Ultra
+3
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
Software Engineer II
Software Engineer II

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 73,000 - 145,000
MacBook Pro
Premium AI tools
Cursor Ultra
+6
AI Evaluation Engineer: Build Verification & Quality Gates
AI Evaluation Engineer: Build Verification & Quality Gates

Zof AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
MacBook Pro
Claude Code Ultra
Research Engineer - Evals
Research Engineer - Evals

Pantera Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash compensation
Equity opportunities
Relocation support
+1
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Seattle (WA)

On-site
USD 230,000 - 385,000
Software Engineer - Evals
Software Engineer - Evals

SpaceXAI • Palo Alto (CA)

On-site
USD 175,000 - 275,000
Equity
Medical coverage
Vision insurance
+5
AI Engineer, Evaluation
AI Engineer, Evaluation

Distyl AI • New York (NY)

Hybrid
USD 150,000 - 250,000
Equity
Medical insurance
Flexible time off
+1
Content Marketing Manager, Technical
Content Marketing Manager, Technical

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 4,000 - 8,000
MacBook Pro
premium AI development tools
Cursor Ultra
+8
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000