The opportunity
We are seeking a highly experienced and detailoriented AI Model Evaluation Consultant with strong expertise in working alongside AI/ML models, Generative AI (GenAI) systems, and multiagent workflows within GxPregulated and qualitycritical environments. The ideal candidate will ensure the reliability, fairness, robustness, and regulatory compliance of AI solutions while engaging directly with crossfunctional teams in designing, testing, and maintaining productionready, inspectionready AI systems.
This role is responsible for leading AI model consultation and development activities across the full model lifecycle, including model integration, training, testing, deployment, and postdeployment monitoring, in alignment with GxP principles, riskbased validation, and data integrity expectations. The consultant will play a key role in embedding Responsible AI, qualitybydesign, and governancebydesign practices, enabling the safe, trustworthy, and explainable use of AI and agentbased solutions in regulated life sciences and pharmaceutical contexts.
Your key responsibilities
- Work on development of AI/ML and Agentic AI strategies by providing analysis, validation insights, and readiness assessments.
- Perform AI/ML model integration, evaluation and validation across development, training, testing, and deployment phases.
- Evaluate model performance using precision, recall, accuracy, F1score, and confusion matrices.
- Gather and analyze business requirements and translate them into functional and technical solution designs.
- Configure, customize, and develop applications, workflows, data models, integrations, and reports based on business needs.
- Assess model behavior across datasets, edge cases, and failure scenarios to identify bias, drift, and instability.
- Demonstrate strong understanding of supervised and unsupervised learning models, including classification, regression, clustering, and anomaly detection.
- Conduct handson evaluation of GenAI and LLMbased systems, including prompt behavior, response quality, and consistency.
- Design and execute testing strategies for AI agents and multiagent workflows, covering orchestration logic, tool invocation, and output validation.
- Hands-on experience with APIs, scripting/programming, and system development lifecycle (SDLC).
- Use LangChain, LangGraph, and Langfuse for agent workflows, observability, traceability, and evaluation logging.
- Collaborate with data scientists and engineers to identify model improvement opportunities and risk mitigations.
- Document evaluation results, limitations, risks, and acceptance criteria in a clear, auditready manner.
Qualifications:
- Bachelors or master’s degree in Life Sciences, Engineering, or related field.
- 5+ years of experience in AI/ML, GenAI, software testing, validation, or quality engineering roles.
- Experience creating user stories, test scripts, validation plans, or AI solution prototypes.
- Exposure to cloud ecosystems such as Azure, AWS, or GCP.
- Consulting or clientfacing experience preferred.
- Excellent communication, documentation, and stakeholder management skills.
Must-Have Skills And Attributes
- Strong understanding of ML models, LLMs, GenAI workflows, agentic systems, and evaluation techniques.
- Strong foundation in machine learning concepts and model evaluation techniques.
- Handson experience with: Precision, Recall, Confusion Matrix, ROC/AUC, Supervised and Unsupervised ML models
- Practical experience with LangChain, LangGraph, and Langfuse.
- Experience evaluating GenAI / LLMbased applications and agents.
- Strong documentation, analytical, and stakeholder communication skills.
- Experience designing test cases, user stories, and acceptance criteria for AI/MLbased systems.
- Ability to validate AI outputs, evaluate accuracy, detect bias, and assess model reliability under varied conditions.
- Handson skills in building prototypes, test harnesses, and demo environments.
- Strong stakeholder management, analytical thinking, and structured communication skills.
- Familiarity with version control, prompt engineering fundamentals, and AIassisted testing tools is preferred.
- Understanding of data integrity principles and electronic records/e-signatures compliance
- Experience in client-facing roles, managing expectations and delivering solutions
- Strong communication and presentation skills
- Ability to troubleshoot application issues and recommend solutions
- Strong teamwork and collaboration mindset
Good-to-Have Skills And Attributes
- Exposure to business development activities and client engagement strategies
- Ability to create innovative insights and contribute to thought leadership
- Understanding of market trends and competitor landscape
- Experience in process optimization and operational efficiency initiatives
- Ability to align technology with business transformation goals
- High attention to detail and analytical thinking
- Adaptability to changing regulatory and business environments
- Strong problem-solving and decision-making abilities