AI Evaluation QA Engineer: Scale Evals & Improve Answers

Socket.dev

New York, Austin, Miami (NY, TX, FL)

On-site

USD 110,000 - 160,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation is seeking a QA / AI Evaluation Engineer to join a highly motivated team. You will run large-scale evaluations, quantify factual grounding, and build a metrics framework that demonstrates quality gains over time.

You will design load tests, automate regression suites, and collaborate with engineering to reproduce and verify fixes. Strong Python, data-science, and LLM evaluation framework experience are essential.

Qualifications

  • Bachelor’s degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science
QA engineering
LLM evaluation
Test automation
Scripting

Education

Bachelor’s Degree in a technical field

Tools

LLM evaluation frameworks
CI/CD tooling

Job description

Appnovation is seeking a QA / AI Evaluation Engineer to join a highly motivated team. You will run large-scale evaluations, quantify factual grounding, and build a metrics framework that demonstrates quality gains over time.

You will design load tests, automate regression suites, and collaborate with engineering to reproduce and verify fixes. Strong Python, data-science, and LLM evaluation framework experience are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation QA Engineer: Scale Testing & Metrics
AI Evaluation QA Engineer: Scale Testing & Metrics

Appnovation • Miami (FL)

On-site
USD 100,000 - 140,000
AI Evaluation QA Engineer — Scale & Improve Answer Quality
AI Evaluation QA Engineer — Scale & Improve Answer Quality

Appnovation • Austin (CO)

On-site
USD 90,000 - 150,000
AI Evaluation Engineer: Scale QA for LLMs
AI Evaluation Engineer: Scale QA for LLMs

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

appnovation • New York (NY)

On-site
USD 120,000 - 160,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • New York (NY), Austin (TX), Miami (FL)

On-site
USD 110,000 - 160,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Miami (FL)

On-site
USD 100,000 - 140,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Austin (CO)

On-site
USD 90,000 - 150,000
AI Evaluation QA Engineer — ML Testing & Validation
AI Evaluation QA Engineer — ML Testing & Validation

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required