QA Engineer for AI Products

Intone Inc

Bellevue (WA)

Hybrid

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Intone Inc. in Bellevue, WA is seeking a QA Engineer specialized in AI products for a 6-month contract. The role is hybrid and requires local candidates for in-person collaboration.

You will design and execute tests for AI/ML features, build automated suites, and evaluate model outputs for accuracy, bias, and safety, while collaborating with data scientists and ML engineers to ensure robust quality across pipelines and deployments.

Qualifications

  • 5+ years of QA and testing experience.
  • Hands-on experience testing LLM and generative AI products.
  • Practical experience with AI/LLM evaluation frameworks such as Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals or TruLens.

Responsibilities

  • Design and execute test plans for AI/ML-driven features, including model outputs and prompts.
  • Build and maintain automated test suites covering functional, regression, integration and API testing.
  • Evaluate model outputs for accuracy, bias, hallucination and edge-case failures.
  • Develop evaluation frameworks and golden datasets to benchmark model performance over time.
  • Test prompt changes, model version upgrades, and fine-tuning outputs for regressions.
  • Perform adversarial and red-team style testing to surface safety and robustness issues.
  • Validate data pipelines feeding into AI models, including data quality and drift detection.
  • Collaborate with data scientists and ML engineers to define acceptance criteria and quality metrics.
  • Test latency, scalability and reliability of AI services under load.
  • Contribute to CI/CD pipelines with automated and model-evaluation tests.
  • Design test cases for LLM-based products such as chatbots and RAG systems.

Skills

QA testing
LLM testing
Python
SQL
Test automation

Tools

Ragas
DeepEval
LangSmith
PromptFoo
OpenAI Evals
TruLens

Job description

This is a 6-month contract position for a QA Engineer specializing in AI products, based in Bellevue, Washington. The role is hybrid and requires local candidates for in-person collaboration.

Responsibilities
  • Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behavior
  • Build and maintain automated test suites covering functional, regression, integration, and API testing
  • Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures
  • Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time
  • Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions
  • Perform adversarial and red-team style testing to surface safety, security, and robustness issues
  • Validate data pipelines feeding into AI models, including data quality, schema, and drift detection
  • Collaborate with data scientists and ML engineers to define acceptance criteria and quality metrics for models
  • Test latency, scalability, and reliability of AI services under load
  • Contribute to CI/CD pipelines, integrating automated and model-evaluation tests
  • Design test cases for LLM-based products such as chatbots, copilots, RAG systems, and AI agents, accounting for non-deterministic and generative outputs
  • Build evaluation suites, scoring rubrics, and golden datasets using AI/LLM evaluation frameworks
  • Validate prompt changes and model/version upgrades against baseline evaluation sets
  • Write evaluation scripts, test harnesses, and data validation logic
Qualifications
  • Required
    • 5+ years of QA and testing experience
    • Hands-on experience testing LLM and generative AI products
    • Practical experience with AI/LLM evaluation frameworks (such as Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, or TruLens)
    • Working knowledge of evaluation metrics for generative AI, including hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, and semantic similarity measures
    • Strong proficiency in Python for writing test scripts and validation logic
    • Experience with SQL and data validation techniques
    • Test automation expertise
  • Preferred
    • Prompt engineering or prompt testing experience
    • CI/CD pipeline automation experience
    • Exposure to Azure, Snowflake, or cloud data platforms
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA Engineer – AI/ML
QA Engineer – AI/ML

Apptad Inc • Frisco (TX)

On-site
USD 120,000 - 160,000
QA Engineer (AI)
QA Engineer (AI)

Apptad Inc • Frisco (TX)

On-site
USD 120,000 - 180,000
Lead II - Software Testing
Lead II - Software Testing

TekWissen LLC • Bellevue (WA)

Hybrid
USD 120,000 - 160,000
QA Engineer
QA Engineer

WebSenor Ltd • United States

On-site
USD 70,000 - 110,000
Application Engineer Senior (7+ Years)
Application Engineer Senior (7+ Years)

Digipulse Technologies, Inc • United States

Remote
USD 120,000 - 170,000
Senior Software QA Engineer
Senior Software QA Engineer

Crimson Education • United States

On-site
USD 80,000 - 110,000
Collaborative work environment
Focus on AI safety and ethics
Opportunities for professional growth
1. AI - QA & Prompt Analyst
1. AI - QA & Prompt Analyst

Trinityaiboston • Boston (MA)

Hybrid
USD 110,000 - 150,000
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
QA Engineer
QA Engineer

RemoteJobsOne • Kentucky

Remote
USD 124,000 - 241,000
AI Software Test Engineer
AI Software Test Engineer

Spectraforce Technologies • Ann Arbor (MI)

On-site
USD 90,000 - 130,000