Senior LLM Evaluation Engineer

Aspire, Jordan

Egypt (PA)

On-site

USD 140,000 - 200,000

Full time

18 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aspire seeks a Senior LLM Evaluation Engineer to define evaluation strategies for Generative AI across our digital platforms. You will design benchmarks, assess factual accuracy and safety, and ensure production readiness of AI-powered applications.

Collaborating with AI Engineers, PMs, MLR, and business stakeholders, you will develop reusable evaluation datasets, document results, and drive prompt and RAG improvements for reliable AI systems.

Qualifications

  • 8+ years of experience in Software Quality, AI Quality Engineering, Machine Learning, Data Science, or related fields.
  • Hands-on experience evaluating Large Language Models (LLMs) or Generative AI applications.
  • Strong understanding of LLM behavior, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Experience identifying hallucinations, reasoning failures, factual inaccuracies, and inconsistent AI outputs.
  • Experience designing evaluation datasets, benchmark scenarios, and acceptance criteria.
  • Strong analytical and communication skills with attention to detail.
  • Experience documenting evaluation results, quality metrics, and production readiness assessments.

Responsibilities

  • Design and execute comprehensive evaluation strategies for LLM-powered applications.
  • Develop benchmark datasets, golden datasets, and evaluation scenarios based on business requirements and approved reference data.
  • Evaluate AI-generated outputs for factual accuracy, consistency, relevance, grounding, and safety.
  • Identify and document AI failures, edge cases, and retrieval issues.
  • Produce detailed evaluation reports, quality scorecards, and release recommendations.
  • Collaborate with AI Engineers to improve prompts, RAG pipelines, embeddings, and model configurations.
  • Validate AI responses against trusted knowledge sources and business rules.

Skills

Generative AI
LLM Evaluation
AI Quality Assurance
Prompt Engineering
RAG
Groundedness Evaluation
Human-in-the-Loop Evaluation
Regression Evaluation
AI Safety & Compliance
API Testing
Python
Evaluation Frameworks

Tools

LangSmith
Ragas
DeepEval
Promptfoo
OpenAI Evals
Playwright

Job description

About the Role

We are seeking an experienced Senior LLM Evaluation Engineer to ensure the quality, accuracy, and reliability of Generative AI solutions used across our digital platforms. In this role, you will define and execute evaluation strategies for Large Language Models (LLMs), ensuring AI-generated content is factually correct, consistent, compliant, and production-ready.

As the quality owner for AI-powered applications, you will design evaluation frameworks, create benchmark datasets, assess model performance, identify hallucinations and reasoning failures, and provide actionable feedback to improve prompts, retrieval pipelines, and model behavior. You will collaborate closely with AI Engineers, Product Managers, Medical, Legal & Regulatory (MLR), and Business stakeholders to ensure AI systems meet the highest quality standards before release.

Key Responsibilities
  • Design and execute comprehensive evaluation strategies for LLM-powered applications.
  • Develop benchmark datasets, golden datasets, and evaluation scenarios based on business requirements and approved reference data.
  • Evaluate AI-generated outputs for:
  • Factual accuracy
  • Consistency
  • Relevance
  • Groundedness
  • Hallucinations
  • Toxicity and safety concerns
  • Perform manual evaluations and human-in-the-loop assessments of AI responses.
  • Design acceptance criteria and quality metrics for Generative AI applications.
  • Execute regression evaluations following model, prompt, RAG, or configuration changes.
  • Identify and document AI failures, edge cases, reasoning errors, prompt failures, and retrieval issues.
  • Produce detailed evaluation reports, quality scorecards, and release recommendations.
  • Collaborate with AI Engineers to improve prompts, RAG pipelines, embeddings, and model configurations based on evaluation findings.
  • Validate AI responses against trusted knowledge sources and business rules.
  • Support continuous improvement by building reusable evaluation datasets and test assets.
  • Define AI quality standards, governance practices, and evaluation methodologies across AI initiatives.
Required Qualifications
  • 8+ years of experience in Software Quality, AI Quality Engineering, Machine Learning, Data Science, or related fields.
  • Hands-on experience evaluating Large Language Models (LLMs) or Generative AI applications.
  • Strong understanding of LLM behavior, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Experience identifying hallucinations, reasoning failures, factual inaccuracies, and inconsistent AI outputs.
  • Experience designing evaluation datasets, benchmark scenarios, and acceptance criteria.
  • Strong analytical and critical thinking skills with exceptional attention to detail.
  • Experience documenting evaluation results, quality metrics, and production readiness assessments.
  • Excellent communication skills with the ability to collaborate across technical and business teams.
Preferred Qualifications
  • Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith Evaluations, Promptfoo, OpenAI Evals, or similar tools.
  • Experience evaluating RAG systems and retrieval quality.
  • Understanding of LLM-as-a-Judge methodologies and human evaluation workflows.
  • Experience testing AI applications in regulated industries such as healthcare, pharmaceutical, legal, or finance.
  • Familiarity with Azure AI Foundry, Azure OpenAI, OpenAI APIs, Anthropic Claude, Gemini, or similar platforms.
  • Experience with Playwright or API testing tools is a plus.
Technical Skills
  • Generative AI
  • LLM Evaluation
  • AI Quality Assurance
  • Prompt Engineering
  • Retrieval-Augmented Generation (RAG)
  • Groundedness Evaluation
  • Human-in-the-Loop Evaluation
  • Regression Evaluation
  • AI Safety & Compliance
  • API Testing
  • Python (preferred)
  • LangSmith, Ragas, DeepEval, Promptfoo, or equivalent evaluation frameworks
Nice to Have
  • Experience with healthcare or regulated AI solutions.
  • Knowledge of AI governance, responsible AI, and model validation practices.
  • Experience working with vector databases and knowledge retrieval systems.
  • Familiarity with automated evaluation pipelines and AI observability platforms.
  • Experience collaborating with cross-functional AI engineering and product teams.
What You'll Bring

You are passionate about ensuring AI systems are trustworthy, accurate, and safe. You understand that building an LLM is only part of the solution—evaluating, validating, and continuously improving its outputs is equally important. You enjoy combining analytical thinking with human judgment to ensure AI applications deliver reliable, production-ready experiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
Large Language Model Engineer - AI & Human Health Research
Large Language Model Engineer - AI & Human Health Research

Mount Sinai Health System • United States

On-site
USD 110,000 - 165,000
Senior LLM Evaluation & Quality Engineer
Senior LLM Evaluation & Quality Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Software Engineer - AI/ML
Senior Software Engineer - AI/ML

Mitratech • Germany (OH)

Hybrid
USD 100,000 - 140,000
Equal opportunity employer policy