AI Model Evaluation Engineer – Build Verifications

Zof AI, Inc.

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

MacBook Pro
Premium AI development tools
OpenAI Codex Max

Job summary

Zof AI, Inc. in San Francisco, CA is seeking a Model Evaluation Engineer to build tests that determine whether AI actually works.

This full-time on-site role sits close to the core product, designing eval suites, verification harnesses, and quality gates to turn performance into measurable evidence. The ideal candidate will be skeptical by default, rigorous about measurement, and comfortable turning “it seems to work” into verifiable data.

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

Testing software systems
LLM/agent failure understanding
Analytical rigor
Automation scripting
Communication
Ownership
Skepticism
Quality assurance

Tools

CI/CD pipelines
Python
Test harnesses
Automation frameworks

Job description

Zof AI, Inc. in San Francisco, CA is seeking a Model Evaluation Engineer to build tests that determine whether AI actually works.

This full-time on-site role sits close to the core product, designing eval suites, verification harnesses, and quality gates to turn performance into measurable evidence. The ideal candidate will be skeptical by default, rigorous about measurement, and comfortable turning “it seems to work” into verifiable data.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer - Build Eval Suites & Quality Gates
AI Evaluation Engineer - Build Eval Suites & Quality Gates

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Model Evaluation Engineer
Model Evaluation Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 190,000
Model Evaluation Engineer
Model Evaluation Engineer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
MacBook Pro
Premium AI development tools
OpenAI Codex Max
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Evaluation Scientist — Real-World ML Research
AI Evaluation Scientist — Real-World ML Research

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Health benefits
+2
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2
AI Model Engineer: Code Review & Evaluation
AI Model Engineer: Code Review & Evaluation

AfterQuery Experts • San Francisco (CA)

On-site
USD 110,000 - 170,000
Competitive Pay
AI Model Evaluation Specialist
AI Model Evaluation Specialist

BAM Ventures • New York (NY)

On-site
USD 100,000 - 130,000
AI QA Automation Engineer — End-to-End Testing + Equity
AI QA Automation Engineer — End-to-End Testing + Equity

Zof AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Equity
MacBook Pro
Premium AI tools
+1
AI Software Engineer — Model Evaluation & Code Review
AI Software Engineer — Model Evaluation & Code Review

AfterQuery Experts • New York (NY)

On-site
USD 90,000 - 130,000
Competitive Pay