About the Role
We are seeking an experienced Senior LLM Evaluation Engineer to ensure the quality, accuracy, and reliability of Generative AI solutions used across our digital platforms. In this role, you will define and execute evaluation strategies for Large Language Models (LLMs), ensuring AI-generated content is factually correct, consistent, compliant, and production-ready.
As the quality owner for AI-powered applications, you will design evaluation frameworks, create benchmark datasets, assess model performance, identify hallucinations and reasoning failures, and provide actionable feedback to improve prompts, retrieval pipelines, and model behavior. You will collaborate closely with AI Engineers, Product Managers, Medical, Legal & Regulatory (MLR), and Business stakeholders to ensure AI systems meet the highest quality standards before release.
Key Responsibilities
- Design and execute comprehensive evaluation strategies for LLM-powered applications.
- Develop benchmark datasets, golden datasets, and evaluation scenarios based on business requirements and approved reference data.
- Evaluate AI-generated outputs for:
- Factual accuracy
- Consistency
- Relevance
- Groundedness
- Hallucinations
- Toxicity and safety concerns
- Perform manual evaluations and human-in-the-loop assessments of AI responses.
- Design acceptance criteria and quality metrics for Generative AI applications.
- Execute regression evaluations following model, prompt, RAG, or configuration changes.
- Identify and document AI failures, edge cases, reasoning errors, prompt failures, and retrieval issues.
- Produce detailed evaluation reports, quality scorecards, and release recommendations.
- Collaborate with AI Engineers to improve prompts, RAG pipelines, embeddings, and model configurations based on evaluation findings.
- Validate AI responses against trusted knowledge sources and business rules.
- Support continuous improvement by building reusable evaluation datasets and test assets.
- Define AI quality standards, governance practices, and evaluation methodologies across AI initiatives.
Required Qualifications
- 8+ years of experience in Software Quality, AI Quality Engineering, Machine Learning, Data Science, or related fields.
- Hands-on experience evaluating Large Language Models (LLMs) or Generative AI applications.
- Strong understanding of LLM behavior, prompt engineering, and Retrieval-Augmented Generation (RAG).
- Experience identifying hallucinations, reasoning failures, factual inaccuracies, and inconsistent AI outputs.
- Experience designing evaluation datasets, benchmark scenarios, and acceptance criteria.
- Strong analytical and critical thinking skills with exceptional attention to detail.
- Experience documenting evaluation results, quality metrics, and production readiness assessments.
- Excellent communication skills with the ability to collaborate across technical and business teams.
Preferred Qualifications
- Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith Evaluations, Promptfoo, OpenAI Evals, or similar tools.
- Experience evaluating RAG systems and retrieval quality.
- Understanding of LLM-as-a-Judge methodologies and human evaluation workflows.
- Experience testing AI applications in regulated industries such as healthcare, pharmaceutical, legal, or finance.
- Familiarity with Azure AI Foundry, Azure OpenAI, OpenAI APIs, Anthropic Claude, Gemini, or similar platforms.
- Experience with Playwright or API testing tools is a plus.
Technical Skills
- Generative AI
- LLM Evaluation
- AI Quality Assurance
- Prompt Engineering
- Retrieval-Augmented Generation (RAG)
- Groundedness Evaluation
- Human-in-the-Loop Evaluation
- Regression Evaluation
- AI Safety & Compliance
- API Testing
- Python (preferred)
- LangSmith, Ragas, DeepEval, Promptfoo, or equivalent evaluation frameworks
Nice to Have
- Experience with healthcare or regulated AI solutions.
- Knowledge of AI governance, responsible AI, and model validation practices.
- Experience working with vector databases and knowledge retrieval systems.
- Familiarity with automated evaluation pipelines and AI observability platforms.
- Experience collaborating with cross-functional AI engineering and product teams.
What You'll Bring
You are passionate about ensuring AI systems are trustworthy, accurate, and safe. You understand that building an LLM is only part of the solution—evaluating, validating, and continuously improving its outputs is equally important. You enjoy combining analytical thinking with human judgment to ensure AI applications deliver reliable, production-ready experiences.