Model Evaluation Engineer

Zof AI, Inc.

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

46 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

MacBook Pro
Premium AI development tools
OpenAI Codex Max

Job summary

Zof AI, Inc. in San Francisco, CA is seeking a Model Evaluation Engineer to build tests that determine whether AI actually works.

This full-time on-site role sits close to the core product, designing eval suites, verification harnesses, and quality gates to turn performance into measurable evidence. The ideal candidate will be skeptical by default, rigorous about measurement, and comfortable turning “it seems to work” into verifiable data.

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

Testing software systems
LLM/agent failure understanding
Analytical rigor
Automation scripting
Communication
Ownership
Skepticism
Quality assurance

Tools

CI/CD pipelines
Python
Test harnesses
Automation frameworks

Job description

Model Evaluation Engineer

San Francisco, CAFull-timeMid to SeniorOn-site

Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for candidates who want to contribute directly to the development of ambitious AI products in a high-performance environment.

Compensation

Competitive salary

Plus meaningful equity

About This Role

Zof AI is seeking a Model Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Responsibilities
  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.
Requirements
  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.
Nice to have
  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.
What we provide in San Francisco
  • MacBook Pro
  • Premium AI development tools
  • Cursor Ultra
  • Claude Code Ultra
  • OpenAI Codex Max or equivalent advanced AI tooling
  • Access to a high-performance AI product environment
  • Close collaboration with leadership, engineering, and customers
  • Opportunity to work in the San Francisco AI ecosystem
  • Wellness and productivity support where applicable
  • Competitive startup environment
  • High ownership
  • Direct product impact

Benefits may depend on role and final offer terms.

EvalsVerificationAI QualityTestingQA

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Evaluation Engineer
Model Evaluation Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 190,000
Generative AI Engineer
Generative AI Engineer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 230,000
MacBook Pro
Premium AI development tools
Cursor Ultra
+3
Technical Product Manager
Technical Product Manager

Zof AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
MacBook Pro
Premium AI tools
Cursor Ultra
+9
Applied Data Scientist
Applied Data Scientist

Zof AI, Inc. • San Francisco (CA)

On-site
USD 120,000 - 190,000
MacBook Pro
Premium AI development tools
High ownership
+1
Software Engineer in Test
Software Engineer in Test

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 180,000
MacBook Pro
Premium AI tools
Cursor Ultra
+2
AI Model Evaluation Engineer – Build Verifications
AI Model Evaluation Engineer – Build Verifications

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
MacBook Pro
Premium AI development tools
OpenAI Codex Max
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 84,000 - 167,000
MacBook Pro
Premium AI development tools
High ownership
+2
AI Evaluation Engineer - Build Eval Suites & Quality Gates
AI Evaluation Engineer - Build Eval Suites & Quality Gates

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Interaction Designer
Interaction Designer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
MacBook Pro
AI tools
Cursor Ultra
+5