LLM QA Engineer: AI Data Quality & Regression (Contract)

Blackhornvc

Austin (TX)

Hybrid

USD 83,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Paid time off
Company holidays
401K matching

Job summary

Sustainment is seeking a QA Engineer to ensure reliability and robustness of AI Agents in an AWS-based setup. You will design automated QA systems, evaluate ground-truth datasets, and score structured extractions from multi-page documents, adapting to changing schemas and production failures.

Candidate should have 3+ years in software QA with ML/NLP focus, strong Python, PyTest, and data-validation skills, and experience with LLM eval and model monitoring.

Qualifications

  • 3+ years in software testing and quality assurance.
  • 2+ years with a focus on ML evaluation, NLP, LLMs, VLMs, etc.
  • Deep understanding of LLM data quality challenges and common failure modes.

Responsibilities

  • Design and run regression test suites for LLM evaluation.
  • Identify and track LLM failure modes, including hallucinations, biases, factual inconsistencies, and logical errors.
  • Design data-quality checks to assess training and test datasets.
  • Automate LLM performance monitoring using advanced metrics and validation strategies.
  • Apply best practices for prompt-engineering testing, fine-tuning validation, and output-consistency analysis.
  • Collaborate with ML engineers, data scientists, and product teams to align on quality benchmarks.
  • Work within an AWS ecosystem, leveraging services such as EKS, S3, SageMaker, or Databricks for model testing and evaluation.
  • Build tools and dashboards to track LLM quality over time.
  • Curate and version the ground-truth datasets that serve as the accuracy baseline for document parsing, and translate business and domain requirements into written, testable field definitions.
  • Evaluate structured extraction from real business documents by scoring model output field-by-field against ground truth, with tolerance-aware comparison for numbers, dates, free text, and repeated structures.
  • Maintain the ground-truth corpus as a versioned, evolving test asset: keep existing annotations valid as extraction schemas change, preserve dataset provenance, and grow the corpus from real production failures so every customer-reported miss becomes a permanent regression case.
  • Calibrate and validate automated scoring itself; confirm that semantic/LLM-judge scoring agrees with human judgment.

Skills

Python
PyTest
Hypothesis
LangSmith
MLflow
SageMaker
Datadog
Kubernetes
Tilt
OpenAI APIs
Anthropic APIs
Bedrock
.NET
EF Core
MLOps
OCR/Document AI

Tools

Kubernetes
Tilt
Datadog
SageMaker
Databricks
EF Core

Job description

Sustainment is seeking a QA Engineer to ensure reliability and robustness of AI Agents in an AWS-based setup. You will design automated QA systems, evaluate ground-truth datasets, and score structured extractions from multi-page documents, adapting to changing schemas and production failures.

Candidate should have 3+ years in software QA with ML/NLP focus, strong Python, PyTest, and data-validation skills, and experience with LLM eval and model monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM QA Engineer – AI Data Quality & Testing
LLM QA Engineer – AI Data Quality & Testing

Sustainment • Austin (TX)

On-site
USD 110,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
QA Engineer - Gen AI
QA Engineer - Gen AI

Blackhornvc • Austin (TX)

Hybrid
USD 83,000 - 165,000
Medical benefits
Dental benefits
Vision benefits
+3
QA Engineer - Gen AI
QA Engineer - Gen AI

Sustainment • Austin (TX)

On-site
USD 110,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
QA Engineer - Gen AI
QA Engineer - Gen AI

Sustainment Technologies Inc. • Austin (TX)

On-site
USD 85,000 - 140,000
Medical
Dental
Vision
+3
AI Quality Engineer: Evaluation Pipelines & LLM Testing
AI Quality Engineer: Evaluation Pipelines & LLM Testing

Dealstitch LLC. • United States

On-site
USD 160,000 - 220,000
Lead AI QA Engineer for LLM & AI Testing
Lead AI QA Engineer for LLM & AI Testing

Jobtailor • Hoboken (NJ)

On-site
USD 120,000 - 160,000
Staff Quality Engineer - AI/LLM Testing & Automation
Staff Quality Engineer - AI/LLM Testing & Automation

Talener • Hoboken (NJ)

Hybrid
USD 130,000 - 150,000
Senior AI Engineer: Production LLM Agents
Senior AI Engineer: Production LLM Agents

Ataccama • Prague (OK)

On-site
USD 140,000 - 190,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
Generative AI QA Engineer | LLM Evaluation & E2E Testing
Generative AI QA Engineer | LLM Evaluation & E2E Testing

Expedite Talent Solutions • United States

On-site
USD 110,000 - 160,000