AI QA & Evaluation Engineer - Scale & Metrics

Appnovation Technologies

New York (NY)

On-site

USD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation Technologies in New York is seeking a QA / AI Evaluation Engineer to join a forward-leaning team. You will run evaluations at scale, measure factual grounding and accuracy lift, and build metrics that show improvements over time.

You'll design load tests, maintain test plans, and automate regression and evaluation suites, integrating them into CI/CD. The role requires strong Python, data-science techniques, experience with LLM evaluation frameworks, and the ability to translate

Qualifications

  • Bachelor’s degree or equivalent in a technical field.
  • 4+ years in QA/test engineering with data/ML exposure.
  • Strong Python and data-science techniques for evaluating factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses and lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Familiarity with test automation frameworks and scripting.
  • Detail-oriented with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvements and clear reporting.
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
QA Testing
Data Science
LLM Evaluation
Test Automation
Scripting
Analytical Skills
Communication

Education

Bachelor’s Degree in a technical field

Job description

Appnovation Technologies in New York is seeking a QA / AI Evaluation Engineer to join a forward-leaning team. You will run evaluations at scale, measure factual grounding and accuracy lift, and build metrics that show improvements over time.

You'll design load tests, maintain test plans, and automate regression and evaluation suites, integrating them into CI/CD. The role requires strong Python, data-science techniques, experience with LLM evaluation frameworks, and the ability to translate

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Scale QA & Metrics
AI Evaluation Engineer: Scale QA & Metrics

Appnovation Technologies • Phoenix (AZ)

On-site
USD 100,000 - 160,000
AI Evaluation QA Engineer: Scale Testing & Metrics
AI Evaluation QA Engineer: Scale Testing & Metrics

Appnovation • Miami (FL)

On-site
USD 100,000 - 140,000
AI Evaluation Engineer: Scale QA for LLMs
AI Evaluation Engineer: Scale QA for LLMs

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
AI Evaluation QA Engineer — Scale & Improve Answer Quality
AI Evaluation QA Engineer — Scale & Improve Answer Quality

Appnovation • Austin (CO)

On-site
USD 90,000 - 150,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • New York (NY)

On-site
USD 90,000 - 150,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Phoenix (AZ)

On-site
USD 100,000 - 160,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Austin (CO)

On-site
USD 90,000 - 150,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Miami (FL)

On-site
USD 100,000 - 140,000
Senior AI Test & Evaluation Engineer
Senior AI Test & Evaluation Engineer

Motion • Birmingham (AL), Northern (KY)

Hybrid
USD 120,000 - 180,000