Stand out for this role — generate a tailored resume and cover letter in about a minute.
Apptad Inc. is seeking an AI test engineer to design and execute test plans for AI/ML features, including model outputs, prompts, and integrated app behavior.
You will build automated test suites for functional, regression, integration, and API testing, and collaborate with data scientists to define acceptance criteria and quality metrics. You will evaluate model outputs for accuracy, bias, and edge cases, develop evaluation datasets, and perform adversarial testing to surface safety and
Job Description
Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behavior
Build and maintain automated test suites covering functional, regression, integration, and API testing
Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures
Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time
Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions
Perform adversarial and red-team style testing to surface safety, security, and robustness issues
Validate data pipelines feeding into AI models (data quality, schema, drift detection)
Collaborate with data scientists/ML engineers to define acceptance criteria and quality metrics for models
Test latency, scalability, and reliability of AI services under load
Contribute to CI/CD pipelines, integrating automated and model-evaluation tests
Hands-on experience testing LLM-based products (chatbots, copilots, RAG systems, AI agents) - designing test cases for non-deterministic, generative outputs
Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) - building eval suites, scoring rubrics, and golden datasets
Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable
Experience with prompt regression testing - validating prompt changes and model/version upgrades against baseline eval sets
Strong proficiency in Python for writing eval scripts, test harnesses, and data validation logic
Familiarity with SQL and data validation techniques
The environment is primarily Microsoft Azure-based, with Snowflake also playing a key role. Their GenAI initiatives largely leverage OpenAI and Claude models.
Must have skills:
LLM/GenAI product testing
Python expertise
AI evaluation frameworks
SQL
Test automation