Model Evaluation Engineer

Zof AI

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zof AI in San Francisco is seeking a Model Evaluation Engineer to design eval suites, verification harnesses, and quality gates for AI systems. You will translate customer needs into testable checks and drive measurable evidence that the product behaves as intended.

Focus on identifying failure modes, such as regressions and hallucinations, and wiring evaluations into CI to raise the standard of what "working" means across the company.

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

QA testing
LLM failures
Analytical rigor
Test automation
Attention to detail
Clear communication
Fast-paced environment
Ownership

Job description

Zof AI is seeking a Model Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA

Responsibilities
  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.
Requirements
  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.
Nice to have
  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.

Must understand how to measure whether AI systems actually work, beyond demos

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer - Build Eval Suites & Quality Gates
AI Evaluation Engineer - Build Eval Suites & Quality Gates

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Generative AI Engineer
Generative AI Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Software Engineer in Test
Software Engineer in Test

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior AI Software Engineer
Senior AI Software Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Analytics Engineer
Analytics Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Applied Data Scientist
Applied Data Scientist

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
AI QA Automation Engineer — End-to-End Testing + Equity
AI QA Automation Engineer — End-to-End Testing + Equity

Zof AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Equity
MacBook Pro
Premium AI tools
+1