Senior AI Evaluation Engineer: LLMs & Agent Training

Chegg

United States

Remote

USD 180,000 - 230,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Chegg is building an AI Evaluation Platform to measure model and agent performance across accuracy, reasoning, safety, and pedagogical quality. You will own evaluation frameworks end-to-end, from framework design through deployment and production monitoring, partnering with ML researchers, applied scientists, and product teams to define what good looks like for models and agents.

This senior role focuses on creating robust evaluation pipelines, rubrics, model-graded evaluators, and

Qualifications

  • > 5+ years of software/ML engineering experience in evals, testing, or ML measurement infra.
  • > Direct experience with foundation model evaluation and benchmarking in frontier labs or similar.
  • > End-to-end ownership of evaluation solutions from design to deployment and monitoring.
  • > Experience designing evals for LLMs/AI agents, including benchmark design and human annotation platforms.
  • > Proficient in Python; strong in building production-grade data/ML pipelines.
  • > Familiar with SOTA training/eval lifecycles (SFT, RLHF/RLAIF/DPO) and post-training evals.
  • > Experience with agentic architectures and tooling for tracing, versioning, and experiment tracking.
  • > Excellent cross-functional collaboration and critical thinking about metric validity.

Responsibilities

  • > Own evaluation solutions end-to-end: design methodology, build pipelines, deploy into training/production, iterate with minimal hand-off.
  • > Design and build eval frameworks for LLMs and multi-step agents, including offline benchmarks and online evals.
  • > Define rubrics for open-ended tasks, ensure reliable model judgments and bias testing.
  • > Develop model-graded evaluators and validate against human judgment.
  • > Create agent-specific evaluations: tool use, planning, latency, and cost assessments.
  • > Generate golden datasets, adversarial tests, and regression suites to catch regressions.
  • > Translate eval results into data curation and training signal improvements (SFT/RLHF/RLAIF).
  • > Instrument production data collection for feedback into evaluation/training loops.
  • > Build dashboards for researchers/PMs/leadership to monitor model quality trends.

Skills

Python
PyTorch
TensorFlow
production-grade pipelines
statistical judgment
cross-functional collaboration
ML eval design
LLM evaluation
risk and safety thinking

Tools

AWS SageMaker
Bedrock
Databricks MLflow
Unity Catalog
Delta Lake

Job description

Chegg is building an AI Evaluation Platform to measure model and agent performance across accuracy, reasoning, safety, and pedagogical quality. You will own evaluation frameworks end-to-end, from framework design through deployment and production monitoring, partnering with ML researchers, applied scientists, and product teams to define what good looks like for models and agents.

This senior role focuses on creating robust evaluation pipelines, rubrics, model-graded evaluators, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Systems Engineer - LLM, Evaluation, Equity
Staff AI Systems Engineer - LLM, Evaluation, Equity

Maven • San Jose (CA)

On-site
USD 180,000 - 240,000
Physical Health Benefits
Mental Health Benefits
Emotional Health Benefits
+5
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Senior Software Engineer - Model Training & AI Evals
Senior Software Engineer - Model Training & AI Evals

Chegg • United States

Remote
USD 180,000 - 230,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000
AI Evaluations Architect, Decision Intelligence
AI Evaluations Architect, Decision Intelligence

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 278,000
Medical insurance
Dental coverage
Retirement benefits
+3
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000