AI Evaluation Engineer

Zof AI

San Francisco, Northern (CA, KY)

On-site

USD 120,000 - 160,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zof AI in San Francisco, CA is seeking an AI Evaluation Engineer to design and build eval suites and verification harnesses for AI products.

You will turn customer requirements into testable checks, investigate failure modes like regressions and hallucinations, and integrate evals into CI. The role emphasizes rigorous measurement, clear communication, and ownership in a fast-moving environment.

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

QA testing
Evaluation
SQL/Script coding
CI integration
Analytical thinking
Communication
Ownership
Attention to detail

Tools

Test automation tooling

Job description

Zof AI is seeking an AI Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA

Responsibilities
  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.
Requirements
  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.
Nice to have
  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.

Must understand how to measure whether AI systems actually work, beyond demos

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Build AI Test Suites & Quality Gates
AI Evaluation Engineer: Build AI Test Suites & Quality Gates

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 160,000
Test Automation Engineer
Test Automation Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 110,000 - 160,000
Applied AI Engineer
Applied AI Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI QA Automation Engineer — End-to-End Testing + Equity
AI QA Automation Engineer — End-to-End Testing + Equity

Zof AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Equity
MacBook Pro
Premium AI tools
+1
User Experience Designer
User Experience Designer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 90,000 - 130,000
Senior AI Applications Engineer
Senior AI Applications Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Full Stack Software Engineer
Full Stack Software Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 110,000 - 165,000
Technical Solutions Engineer
Technical Solutions Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 160,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000