QA Engineer - Gen AI

Blackhornvc

Austin (TX)

Hybrid

USD 83,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Paid time off
Company holidays
401K matching

Job summary

Sustainment is seeking a QA Engineer to ensure reliability and robustness of AI Agents in an AWS-based setup. You will design automated QA systems, evaluate ground-truth datasets, and score structured extractions from multi-page documents, adapting to changing schemas and production failures.

Candidate should have 3+ years in software QA with ML/NLP focus, strong Python, PyTest, and data-validation skills, and experience with LLM eval and model monitoring.

Qualifications

  • 3+ years in software testing and quality assurance.
  • 2+ years with a focus on ML evaluation, NLP, LLMs, VLMs, etc.
  • Deep understanding of LLM data quality challenges and common failure modes.

Responsibilities

  • Design and run regression test suites for LLM evaluation.
  • Identify and track LLM failure modes, including hallucinations, biases, factual inconsistencies, and logical errors.
  • Design data-quality checks to assess training and test datasets.
  • Automate LLM performance monitoring using advanced metrics and validation strategies.
  • Apply best practices for prompt-engineering testing, fine-tuning validation, and output-consistency analysis.
  • Collaborate with ML engineers, data scientists, and product teams to align on quality benchmarks.
  • Work within an AWS ecosystem, leveraging services such as EKS, S3, SageMaker, or Databricks for model testing and evaluation.
  • Build tools and dashboards to track LLM quality over time.
  • Curate and version the ground-truth datasets that serve as the accuracy baseline for document parsing, and translate business and domain requirements into written, testable field definitions.
  • Evaluate structured extraction from real business documents by scoring model output field-by-field against ground truth, with tolerance-aware comparison for numbers, dates, free text, and repeated structures.
  • Maintain the ground-truth corpus as a versioned, evolving test asset: keep existing annotations valid as extraction schemas change, preserve dataset provenance, and grow the corpus from real production failures so every customer-reported miss becomes a permanent regression case.
  • Calibrate and validate automated scoring itself; confirm that semantic/LLM-judge scoring agrees with human judgment.

Skills

Python
PyTest
Hypothesis
LangSmith
MLflow
SageMaker
Datadog
Kubernetes
Tilt
OpenAI APIs
Anthropic APIs
Bedrock
.NET
EF Core
MLOps
OCR/Document AI

Tools

Kubernetes
Tilt
Datadog
SageMaker
Databricks
EF Core

Job description

Company Overview:

Sustainment is an AI-native software platform that helps US-based manufacturers easily find and work with the critical suppliers they need to build and manage their supply chains. Our vision is to reimagine American manufacturing as a hyperconnected, secure, and resilient ecosystem of local and regional suppliers who can more easily connect, interact, and do business with the industry and government customers that rely on them. We are a dual-use technology platform that supports both DoD and commercial customers in pursuit of our vision.

This is a contract opportunity
Job Overview:

We are seeking a QA Engineer to help ensure the reliability, accuracy, and robustness of our AI Agents. This role will focus on data quality, model evaluation, and regression testing frameworks to identify and mitigate common LLM failure modes. You will be responsible for designing automated and scalable quality assurance systems while working in an AWS-based infrastructure. If you have a strong background in LLM testing, data validation, and automated QA frameworks, this role is an excellent opportunity to contribute to cutting-edge AI systems.

Responsibilities:
  • Design and run regression test suites for LLM evaluation.
  • Identify and track LLM failure modes, including hallucinations, biases, factual inconsistencies, and logical errors.
  • Design data-quality checks to assess training and test datasets.
  • Automate LLM performance monitoring using advanced metrics and validation strategies.
  • Apply best practices for prompt-engineering testing, fine-tuning validation, and output-consistency analysis.
  • Collaborate with ML engineers, data scientists, and product teams to align on quality benchmarks.
  • Work within an AWS ecosystem, leveraging services such as EKS, S3, SageMaker, or Databricks for model testing and evaluation.
  • Build tools and dashboards to track LLM quality over time.
  • Curate and version the ground-truth datasets that serve as the accuracy baseline for document parsing, and translate business and domain requirements into written, testable field definitions (partnering with the labeling team on annotation guidelines).
  • Evaluate structured extraction from real business documents (multi-page PDFs, scans, spreadsheets) by scoring model output field-by-field against ground truth, with tolerance-aware comparison for numbers, dates, free text, and repeated structures.
  • Maintain the ground-truth corpus as a versioned, evolving test asset: keep existing annotations valid as extraction schemas change, preserve dataset provenance, and grow the corpus from real production failures so every customer-reported miss becomes a permanent regression case.
  • Calibrate and validate automated scoring itself; confirm that semantic/LLM-judge scoring agrees with human judgment.
Qualifications:
  • 3+ years in software testing and quality assurance
  • 2+ years with a focus on ML evaluation, NLP, LLMs, VLMs, etc.
  • Deep understanding of LLM data quality challenges and common failure modes.
  • Experience designing automated tests for AI/ML models.
  • Familiarity with Python and testing frameworks such as PyTest, Hypothesis, or similar.
  • Knowledge of evaluation metrics for LLMs (DeepEval, MLflow, LangSmith, or similar).
  • Hands-on experience with automated data validation techniques.
  • Strong debugging and analytical skills.
  • Experience creating or working with labeled evaluation datasets (“golden” sets) for model evaluation.
  • Working knowledge of evaluation metrics for structured information extraction: field-level precision, recall, and F1; exact vs. fuzzy matching; numeric tolerance; and alignment of repeated or nested records.
  • Experience translating ambiguous business requirements into precise, documented field definitions in collaboration with non-technical subject-matter experts.
Preferred Qualifications
  • SQL proficiency, including seeding test data across Postgres environments (local/dev/staging/prod).
  • Comfort with observability and incident-response tooling (e.g., Datadog monitors, alerting/triage) for monitoring and debugging.
  • Familiarity with RAG and RAGAS.
  • Familiarity with containerized dev environments (Kubernetes/Tilt).
  • Understanding of human-in-the-loop (HITL) evaluation strategies.
  • Familiarity with LLM APIs (OpenAI, Anthropic, Bedrock, or similar).
  • Background in statistical analysis or model interpretability.
  • Experience with MLOps practices and CI/CD pipelines for ML models.
  • Experience evaluating document AI / OCR pipelines and their specific failure modes: layout and table extraction, multi-page documents, scanned or low-quality source material.
  • Experience running controlled models and prompt comparison studies.
  • Familiarity with the .NET+Linux ecosystem.
  • Able to read DB schema changes & migrations (EF Core/.NET, DDL).

Sustainment offers a competitive benefits package for full time employees including medical, dental, vision, paid time off, company holidays, and 401K matching.

Sustainment is proud to be an equal opportunity employer. We provide employment opportunities without regard to age, race, color, ancestry, national origin, religion, disability, sex, gender identity or expression, sexual orientation, veteran status, or any other protected class.

Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time.

Sustainment participates in E-Verify.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

QA Engineer - Gen AI New Austin, Texas, United States
QA Engineer - Gen AI New Austin, Texas, United States

Sustainment Technologies Inc. • Austin (TX)

On-site
USD 90,000 - 130,000
Medical, dental, vision coverage
Paid time off
401K matching
QA Engineer - Gen AI
QA Engineer - Gen AI

Sustainment • Austin (TX)

On-site
USD 110,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
QA Engineer - Gen AI
QA Engineer - Gen AI

Sustainment Technologies Inc. • Austin (TX)

On-site
USD 85,000 - 140,000
Medical
Dental
Vision
+3
LLM QA Engineer: AI Data Quality & Regression (Contract)
LLM QA Engineer: AI Data Quality & Regression (Contract)

Blackhornvc • Austin (TX)

Hybrid
USD 83,000 - 165,000
Medical benefits
Dental benefits
Vision benefits
+3
LLM QA Engineer – AI Data Quality & Testing
LLM QA Engineer – AI Data Quality & Testing

Sustainment • Austin (TX)

On-site
USD 110,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
AI/LLM QA Engineer - Data Quality & Validation
AI/LLM QA Engineer - Data Quality & Validation

Sustainment Technologies Inc. • Austin (TX)

On-site
USD 90,000 - 130,000
Medical, dental, vision coverage
Paid time off
401K matching
QA Engineer
QA Engineer

DataServ • Massachusetts

Hybrid
USD 70,000 - 90,000
401(k) with match
Medical, dental, vision benefits
Paid holidays and PTO
+1
QA Engineer
QA Engineer

Dataservtech • Andover (MA)

On-site
USD 85,000 - 110,000
Paid holidays
PTO
Discretionary bonuses
+7
AI Engineer
AI Engineer

Que Technology Group • Fort Meade (MD)

On-site
USD 120,000 - 160,000
Health benefits
Disability insurance
Annual leave 25 days
+2
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 210,000