Job Description
ICE's AI Center of Excellence is looking for a Senior Software Test Engineer to own QA for our AI powered applications. This role works with Systems Analysts, Development, and QA to understand business/product and system requirements and define test scenarios for AI applications, combining core QA fundamentals, Python scripting, and API validation with specialized AI/LLM evaluation skills such as detecting hallucinations, bias, and drift.
Responsibilities
- Lead test strategy for AI/ML features: model output validation, regression testing across model/prompt versions, and edge case discovery.
- Build evaluation frameworks for testing AI outputs for accuracy, consistency, bias, hallucination, and safety/guardrail compliance.
- Define golden datasets, ground truth labels, and release acceptance criteria with data science and ML engineering.
- Automate prompt based test suites and model evaluation metrics into CI/CD release gates.
- Lead and mentor a QA team, setting AI testing standards across projects.
- Translate AI feature requirements into test scenarios and quality gates.
- Monitor production of AI behavior and evolve test coverage based on real world failures.
Knowledge and Experience
- 8+ years of QA/test engineering experience, including 3+ years testing AI, ML, or data driven applications.
- 3+ years of experience with AI/LLM evaluation frameworks: prompt regression, hallucination detection, bias/fairness testing, and model/data drift detection.
- 5+ years of Python experience for test automation, data analysis, and evaluation scripting.
- 3+ years of experience with UI automation frameworks (Selenium, Playwright, or Cypress) for AI powered features.
- Strong SQL and data validation skills for testing the pipelines that feed AI models.
- Working knowledge of ML fundamentals: evaluation of metrics, train/validation/test splits, model versioning.
- Experience testing conversational AI/chatbots: multi turn dialogue, intent/context handling, and guardrail validation against harmful or adversarial prompts.
- Experience testing REST/GraphQL APIs: functional, contract, and schema validation, auth testing, negative/boundary testing, and load testing.
- Experience testing AI applications for infosec: prompt injection resistance, data leakage/PII exposure prevention, access control, and secure endpoint authentication, per organizational security policy.
- Ability to validate AI applications against organizational compliance requirements and frameworks (SOC 2, ISO 27001, NIST AI RMF, GDPR), coordinating with InfoSec and Compliance.
Preferred Knowledge and Experience
- Experience testing RAG systems or autonomous agents.
- Familiarity with responsible AI principles: fairness, transparency, safety guardrails.
- Experience with LLM eval tools such as LangSmith, Ragas, or DeepEval.
- Exposure cloud AI/ML platforms (AWS SageMaker, Azure ML, or GCP Vertex AI).
- Experience in financial services, Mortgage or fintech applications.
- Experience with Git or other version control systems.