AI Evaluation Engineer: Scale QA for LLMs

Appnovation

Dallas (TX)

On-site

USD 95,000 - 140,000

Full time

28 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation is seeking a QA / AI Evaluation Engineer to join a high‑performing team in a forward‑leaning role. You will run large‑scale evaluations to measure factual grounding and accuracy lift, building metrics and test plans that drive quality improvements across our platform.

You will work with engineering to reproduce fixes, design load tests, and automate regression suites, while communicating results clearly to both technical and non‑technical stakeholders.

Qualifications

  • Bachelor’s Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers' outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., 'answer is X% supported by source content / Y% better').
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

QA experience
Python programming
Statistical analysis
LLM evaluation
CI/CD tooling
Communication skills
Problem solving

Education

Bachelor’s degree in a technical field

Tools

Python
LLM evaluation frameworks
CI/CD

Job description

Appnovation is seeking a QA / AI Evaluation Engineer to join a high‑performing team in a forward‑leaning role. You will run large‑scale evaluations to measure factual grounding and accuracy lift, building metrics and test plans that drive quality improvements across our platform.

You will work with engineering to reproduce fixes, design load tests, and automate regression suites, while communicating results clearly to both technical and non‑technical stakeholders.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation QA Engineer — Scale, Metrics & Automation
AI Evaluation QA Engineer — Scale, Metrics & Automation

Appnovation • New York (NY)

On-site
USD 90,000 - 130,000
AI Evaluation QA Engineer: Scale Testing & Metrics
AI Evaluation QA Engineer: Scale Testing & Metrics

Appnovation • Miami (FL)

On-site
USD 100,000 - 140,000
AI Evaluation QA Engineer — Scale & Improve Answer Quality
AI Evaluation QA Engineer — Scale & Improve Answer Quality

Appnovation • Austin (CO)

On-site
USD 90,000 - 150,000
Senior AI Quality Engineer — LLMs & Evaluation Systems
Senior AI Quality Engineer — LLMs & Evaluation Systems

Block • San Francisco (CA)

On-site
USD 190,000 - 230,000
AI Evaluation QA Engineer — ML Testing & Validation
AI Evaluation QA Engineer — ML Testing & Validation

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
Senior LLM Evaluation & Quality Engineer
Senior LLM Evaluation & Quality Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
Generative AI QA Engineer | LLM Evaluation & E2E Testing
Generative AI QA Engineer | LLM Evaluation & E2E Testing

Expedite Talent Solutions • United States

On-site
USD 110,000 - 160,000
Lead AI QA Engineer for LLM & AI Testing
Lead AI QA Engineer for LLM & AI Testing

Jobtailor • Hoboken (NJ)

On-site
USD 120,000 - 160,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000