Prudential is scaling the adoption of AI across the Group, including internally developed AI use cases, vendor solutions and off-the-shelf AI platforms. As more AI systems move into production, we need a strong AI testing and assurance capability to ensure these solutions perform as expected, manage risk appropriately and deliver trusted business outcomes.
The Head of AI Testing & Assurance will be responsible for designing and executing testing approaches across AI and machine learning solutions, including systems built on platforms such as Databricks, GCP / Vertex AI and enterprise AI environments. The role will also support independent third-party assessments where required, but the primary focus will be on hands-on testing, technical validation and assurance.
This is a highly technical hands on role suited to someone with strong experience in AI / ML testing, model evaluation, AI assurance, model validation, MLOps / AI Ops and production AI environments. The individual should be comfortable moving between hands-on technical evaluation and senior stakeholder discussions.
Responsibilities:
AI Testing, Validation and Assurance
- Design and execute comprehensive testing strategies for AI systems, including traditional machine learning models, Generative AI applications and Agentic AI workflows.
- Develop practical test plans, test cases, test datasets and evaluation approaches to assess AI model performance, reliability, safety, fairness and business suitability.
- Conduct hands-on AI model testing, including accuracy testing, robustness testing, bias and fairness testing, explainability review, edge-case testing and reliability assessment.
- Perform technical evaluations of internally developed AI use cases as well as vendor-provided and off-the-shelf AI solutions.
- Build and execute AI assurance checks to ensure AI systems are fit for deployment and continue to perform appropriately in production environments.
- Assess AI solutions against key dimensions such as performance, stability, fairness, explainability, security, scalability and operational resilience.
Generative AI and Agentic AI Testing
- Develop testing approaches for Generative AI solutions, including LLM-based applications, copilots, chatbots, summarisation tools, document intelligence solutions and knowledge retrieval systems.
- Design and execute LLM evaluation test cases covering hallucination risk, factual accuracy, relevance, response quality, prompt sensitivity, toxicity, bias and safety.
- Conduct adversarial testing, jailbreak testing, prompt injection testing and red-team style assessments for high-risk GenAI use cases.
- Test Agentic AI systems for task completion, tool usage, decision logic, workflow reliability, guardrail effectiveness and failure handling.
- Support the development of reusable evaluation frameworks for GenAI and Agentic AI solutions across the Group.
Automated AI Evaluation and Testing Frameworks
- Build a library of reusable AI tests and evaluation templates for common use cases such as OCR, recommendation models, cross-sell models, document processing, customer servicing tools and GenAI applications.
- Develop automated testing and validation checkpoints that can be embedded into CI/CD pipelines, SDLC processes and MLOps / AI Ops workflows.
- Create modular evaluation tools and testing frameworks that can be reused across Traditional AI, Generative AI and Agentic AI systems.
- Support model monitoring activities by developing tests for model drift, performance degradation, data quality issues and changes in model behaviour.
- Work with engineering and platform teams to integrate AI testing into production deployment processes.
AI Risk, Governance and Responsible AI Support
- Support the design of risk-based testing approaches, where higher-risk AI systems receive deeper validation and more rigorous assurance.
- Contribute to AI governance processes by providing technical testing evidence, validation outcomes and assurance recommendations.
- Ensure AI systems are tested against responsible AI principles, including fairness, transparency, robustness, accountability and reliability.
- Provide technical input into AI risk assessments, model approval processes and post-deployment monitoring.
- Support independent third-party assessments where required, including preparing technical documentation, test evidence and validation results.
Requirements:
- Strong hands-on experience in AI testing, AI assurance, model validation, machine learning evaluation or responsible AI assessment.
- Solid understanding of AI and machine learning evaluation methods, including performance testing, robustness testing, fairness testing, explainability and model reliability.
- Experience testing AI or machine learning solutions in production or near-production environments.
- Experience designing and executing test frameworks for Traditional AI, Generative AI and / or Agentic AI systems.
- Strong technical understanding of MLOps, AI Ops, model lifecycle management, CI/CD and production AI deployment.
- Experience with platforms such as Databricks, Azure Databricks, Google Cloud / Vertex AI, MLflow or similar AI / ML platforms.
- Ability to develop reusable test libraries, evaluation tools, validation datasets and automated testing checkpoints.
- Experience with API testing, data pipelines, model monitoring and post-deployment validation.
- Strong communication skills, with the ability to explain technical testing results clearly to business, technology and risk stakeholders.