QA Engineer for AI Products / Solutions
Location: Bellevue, WA & Frisco, TX
Job Overview
We are seeking an experienced QA Engineer to support quality engineering for AI/ML and Generative AI products and solutions. The ideal candidate will have strong expertise in Python, test automation, LLM/GenAI testing, AI evaluation frameworks, and SQL, with hands-on experience validating non-deterministic AI outputs.
Key Responsibilities
- Design and execute comprehensive test plans for AI/ML-driven features, including model outputs, prompts, APIs, and integrated application behavior.
- Build and maintain automated test suites covering functional, regression, integration, and API testing.
- Test LLM-based products including chatbots, copilots, RAG applications, and AI agents.
- Evaluate AI model outputs for accuracy, consistency, relevance, hallucinations, bias, toxicity, and edge-case failures.
- Develop evaluation frameworks, golden datasets, test cases, and scoring methodologies to benchmark model performance.
- Perform prompt regression testing and validate model/version upgrades and fine-tuning changes against established baselines.
- Conduct adversarial and red-team testing to identify safety, security, robustness, and reliability issues.
- Validate data pipelines supporting AI models, including data quality, schema validation, and data drift detection.
- Collaborate with Data Scientists and ML Engineers to establish acceptance criteria, quality metrics, and evaluation strategies.
- Test AI services for latency, scalability, reliability, and performance under load.
- Integrate automated functional and AI evaluation tests into CI/CD pipelines.
- Develop Python-based test harnesses, evaluation scripts, and data validation utilities.
- Use SQL for data validation, test-data analysis, and backend verification.
Required Skills
- Strong hands-on experience with LLM / Generative AI product testing
- Strong Python programming skills
- Experience with AI/LLM evaluation frameworks, such as:
- Ragas
- DeepEval
- LangSmith
- Promptfoo
- OpenAI Evals
- TruLens
- Strong test automation experience
- Strong SQL and data validation skills
- Experience testing RAG, LLM applications, chatbots, copilots, or AI agents
- Understanding of GenAI evaluation metrics such as:
- Hallucination rate
- Faithfulness / Groundedness
- Relevance
- Answer correctness
- Toxicity / Bias
- Semantic similarity
- BLEU / ROUGE where applicable
Preferred / Nice-to-Have Skills
- Prompt engineering and prompt testing
- Prompt regression testing
- CI/CD pipeline automation
- Azure experience
- Snowflake or other cloud data platforms
- Experience with AI/ML data pipelines
- Adversarial or red-team testing of AI applications
- Performance and load testing of AI services