AI Evaluation Engineer - Build Eval Suites & Quality Gates

Zof AI

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Zof AI in San Francisco is seeking a Model Evaluation Engineer to design eval suites, verification harnesses, and quality gates for AI systems. You will translate customer needs into testable checks and drive measurable evidence that the product behaves as intended.

Focus on identifying failure modes, such as regressions and hallucinations, and wiring evaluations into CI to raise the standard of what "working" means across the company.

Qualifications

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Skills

QA testing
LLM failures
Analytical rigor
Test automation
Attention to detail
Clear communication
Fast-paced environment
Ownership

Job description

Zof AI in San Francisco is seeking a Model Evaluation Engineer to design eval suites, verification harnesses, and quality gates for AI systems. You will translate customer needs into testable checks and drive measurable evidence that the product behaves as intended.

Focus on identifying failure modes, such as regressions and hallucinations, and wiring evaluations into CI to raise the standard of what "working" means across the company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Evaluation Engineer
Model Evaluation Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI QA Automation Engineer — End-to-End Testing + Equity
AI QA Automation Engineer — End-to-End Testing + Equity

Zof AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Equity
MacBook Pro
Premium AI tools
+1
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Delivery Engineer: AI Safety & Enterprise Evaluations
Delivery Engineer: AI Safety & Enterprise Evaluations

Artificial Intelligence Underwriting Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive salary
Equity
Relocation to San Francisco
+1
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Alignment Research Engineer — Evaluation & Safety
AI Alignment Research Engineer — Evaluation & Safety

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2
AI Engineer, Evals & Agent Quality
AI Engineer, Evals & Agent Quality

Town.com, Inc. • San Francisco (CA)

On-site
USD 250,000 - 300,000
AI Quality Analyst: Model Evaluation & Safety
AI Quality Analyst: Model Evaluation & Safety

RXinsider LTD. • Lincolnshire (IL)

Hybrid
USD 123,000 - 184,000
Hybrid work
Healthcare
Well-being day
+1
QA Engineer for AI Evaluation & Quality Assurance
QA Engineer for AI Evaluation & Quality Assurance

EVB • United States

Remote
USD 90,000 - 130,000