Senior LLM Evaluation & Quality Engineer

Aspire, Jordan

Egypt (PA)

On-site

USD 140,000 - 200,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aspire seeks a Senior LLM Evaluation Engineer to define evaluation strategies for Generative AI across our digital platforms. You will design benchmarks, assess factual accuracy and safety, and ensure production readiness of AI-powered applications.

Collaborating with AI Engineers, PMs, MLR, and business stakeholders, you will develop reusable evaluation datasets, document results, and drive prompt and RAG improvements for reliable AI systems.

Qualifications

  • 8+ years of experience in Software Quality, AI Quality Engineering, Machine Learning, Data Science, or related fields.
  • Hands-on experience evaluating Large Language Models (LLMs) or Generative AI applications.
  • Strong understanding of LLM behavior, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Experience identifying hallucinations, reasoning failures, factual inaccuracies, and inconsistent AI outputs.
  • Experience designing evaluation datasets, benchmark scenarios, and acceptance criteria.
  • Strong analytical and communication skills with attention to detail.
  • Experience documenting evaluation results, quality metrics, and production readiness assessments.

Responsibilities

  • Design and execute comprehensive evaluation strategies for LLM-powered applications.
  • Develop benchmark datasets, golden datasets, and evaluation scenarios based on business requirements and approved reference data.
  • Evaluate AI-generated outputs for factual accuracy, consistency, relevance, grounding, and safety.
  • Identify and document AI failures, edge cases, and retrieval issues.
  • Produce detailed evaluation reports, quality scorecards, and release recommendations.
  • Collaborate with AI Engineers to improve prompts, RAG pipelines, embeddings, and model configurations.
  • Validate AI responses against trusted knowledge sources and business rules.

Skills

Generative AI
LLM Evaluation
AI Quality Assurance
Prompt Engineering
RAG
Groundedness Evaluation
Human-in-the-Loop Evaluation
Regression Evaluation
AI Safety & Compliance
API Testing
Python
Evaluation Frameworks

Tools

LangSmith
Ragas
DeepEval
Promptfoo
OpenAI Evals
Playwright

Job description

Aspire seeks a Senior LLM Evaluation Engineer to define evaluation strategies for Generative AI across our digital platforms. You will design benchmarks, assess factual accuracy and safety, and ensure production readiness of AI-powered applications.

Collaborating with AI Engineers, PMs, MLR, and business stakeholders, you will develop reusable evaluation datasets, document results, and drive prompt and RAG improvements for reliable AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Senior AI Engineer: LLM Evaluation, Production & Optimization
Senior AI Engineer: LLM Evaluation, Production & Optimization

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
Senior AI Quality Engineer — LLMs & Evaluation Systems
Senior AI Quality Engineer — LLMs & Evaluation Systems

Block • San Francisco (CA)

On-site
USD 190,000 - 230,000
Lead AI Engineer — LLM Evaluation & Production Optimizer
Lead AI Engineer — LLM Evaluation & Production Optimizer

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
Senior ML Evaluation Engineer - LLM Quality & Gates
Senior ML Evaluation Engineer - LLM Quality & Gates

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Senior GenAI & LLM Evaluation Engineer - Remote US
Senior GenAI & LLM Evaluation Engineer - Remote US

Acuity Analytics • United States

On-site
USD 140,000 - 190,000