QA Engineer – AI/ML

Apptad Inc

Frisco (TX)

On-site

USD 120,000 - 160,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apptad Inc. in Bellevue, WA and Frisco, TX is hiring a QA Engineer for AI products. You will design and run QA strategies for AI/ML features, validate non-deterministic outputs, and ensure data pipelines quality.

You will build automated test suites, implement evaluation frameworks, test prompts, APIs, and integrate tests into CI/CD. Strong Python, SQL, and experience with RAG/LLM apps required.

Qualifications

  • Experienced in testing AI/ML features and AI product workflows.
  • Proficient in Python and data validation with SQL.
  • Familiar with AI evaluation metrics and non-deterministic outputs.

Responsibilities

  • Design and execute QA plans for AI/ML features, prompts, APIs, and app behavior.
  • Build and maintain automated test suites for functional, regression, and API testing.
  • Test LLM-based products like chatbots, copilots, and RAG apps.
  • Evaluate model outputs for accuracy, bias, hallucinations, and edge cases.
  • Develop evaluation frameworks, datasets, and scoring methods.
  • Conduct prompt regression tests and validate upgrades against baselines.
  • Perform adversarial testing for safety, security, and reliability.
  • Validate data pipelines, quality, schemas, and drift detection.
  • Collaborate with Data Scientists to set quality metrics and criteria.
  • Test AI services for latency, scalability, and performance under load.
  • Integrate automated tests into CI/CD pipelines.
  • Develop Python test harnesses and data validation utilities.
  • Use SQL for data validation and backend verification.

Skills

LLM testing
Python
Test automation
SQL
AI evaluation
RAG apps
Chatbots
Copilots
AI agents

Tools

Ragas
DeepEval
LangSmith
Promptfoo
OpenAI Evals
TruLens

Job description

QA Engineer for AI Products / Solutions

Location: Bellevue, WA & Frisco, TX

Job Overview

We are seeking an experienced QA Engineer to support quality engineering for AI/ML and Generative AI products and solutions. The ideal candidate will have strong expertise in Python, test automation, LLM/GenAI testing, AI evaluation frameworks, and SQL, with hands-on experience validating non-deterministic AI outputs.

Key Responsibilities
  • Design and execute comprehensive test plans for AI/ML-driven features, including model outputs, prompts, APIs, and integrated application behavior.
  • Build and maintain automated test suites covering functional, regression, integration, and API testing.
  • Test LLM-based products including chatbots, copilots, RAG applications, and AI agents.
  • Evaluate AI model outputs for accuracy, consistency, relevance, hallucinations, bias, toxicity, and edge-case failures.
  • Develop evaluation frameworks, golden datasets, test cases, and scoring methodologies to benchmark model performance.
  • Perform prompt regression testing and validate model/version upgrades and fine-tuning changes against established baselines.
  • Conduct adversarial and red-team testing to identify safety, security, robustness, and reliability issues.
  • Validate data pipelines supporting AI models, including data quality, schema validation, and data drift detection.
  • Collaborate with Data Scientists and ML Engineers to establish acceptance criteria, quality metrics, and evaluation strategies.
  • Test AI services for latency, scalability, reliability, and performance under load.
  • Integrate automated functional and AI evaluation tests into CI/CD pipelines.
  • Develop Python-based test harnesses, evaluation scripts, and data validation utilities.
  • Use SQL for data validation, test-data analysis, and backend verification.
Required Skills
  • Strong hands-on experience with LLM / Generative AI product testing
  • Strong Python programming skills
  • Experience with AI/LLM evaluation frameworks, such as:
    • Ragas
    • DeepEval
    • LangSmith
    • Promptfoo
    • OpenAI Evals
    • TruLens
  • Strong test automation experience
  • Strong SQL and data validation skills
  • Experience testing RAG, LLM applications, chatbots, copilots, or AI agents
  • Understanding of GenAI evaluation metrics such as:
    • Hallucination rate
    • Faithfulness / Groundedness
    • Relevance
    • Answer correctness
    • Toxicity / Bias
    • Semantic similarity
    • BLEU / ROUGE where applicable
Preferred / Nice-to-Have Skills
  • Prompt engineering and prompt testing
  • Prompt regression testing
  • CI/CD pipeline automation
  • Azure experience
  • Snowflake or other cloud data platforms
  • Experience with AI/ML data pipelines
  • Adversarial or red-team testing of AI applications
  • Performance and load testing of AI services
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA Engineer for AI Products
QA Engineer for AI Products

Intone Inc • Bellevue (WA)

Hybrid
USD 120,000 - 160,000
QA Engineer (AI)
QA Engineer (AI)

Apptad Inc • Frisco (TX)

On-site
USD 120,000 - 180,000
QA Engineer
QA Engineer

WebSenor Ltd • United States

On-site
USD 70,000 - 110,000
AI/ML QA Engineer — GenAI & LLM Testing
AI/ML QA Engineer — GenAI & LLM Testing

Apptad Inc • Frisco (TX)

On-site
USD 120,000 - 160,000
Senior Software QA Engineer
Senior Software QA Engineer

Crimson Education • United States

On-site
USD 80,000 - 110,000
Collaborative work environment
Focus on AI safety and ethics
Opportunities for professional growth
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
Application Engineer Senior (7+ Years)
Application Engineer Senior (7+ Years)

Digipulse Technologies, Inc • United States

Remote
USD 120,000 - 170,000
Software Quality Assurance Engineer
Software Quality Assurance Engineer

nexacode • United States

On-site
USD 70,000 - 110,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
AI Software Test Engineer
AI Software Test Engineer

Spectraforce Technologies • Ann Arbor (MI)

On-site
USD 90,000 - 130,000