AI Evaluation Scientist — Metrics, Safety & Trust

UNAVAILABLE

McLean (VA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

UNAVAILABLE is seeking an AI Evaluation Scientist to design and execute evaluation programs for predictive and generative AI systems, ensuring accuracy, safety, and alignment with mission requirements. This role will establish trust in AI solutions and support continuous improvement across the AI lifecycle.

The position collaborates with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support

Qualifications

  • Ability to hold a position of public trust with the U.S. government.
  • Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Data Science, or a related field.
  • 2+ years evaluating ML models, NLP systems, or generative AI models (LLMs preferred).
  • Proficiency in Python and libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Experience in structured error analysis and behavioral audits of AI models.

Responsibilities

  • Implement evaluation frameworks for AI models, including accuracy, robustness, bias, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines to detect drift.
  • Develop benchmark datasets, challenge sets, and scenario-based tests.
  • Perform structured error analysis and audits of LLMs and RAG systems.
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists on experiments and model hardening.
  • Contribute to human-in-the-loop evaluation workflows and governance-aligned reports.
  • Map evaluation outcomes to principles such as fairness, transparency, and safety.
  • Work with AI Governance Analysts to support compliance and risk assessments.
  • Stay current with evaluation tools, metrics, and research in LLM assessment.

Skills

Python
PyTorch
Hugging Face
scikit-learn
LangChain
LLMs
NLP
Evaluation metrics
Experiment design
Data analysis

Education

Bachelor's or Master's in CS/Statistics/ML or related field

Tools

Ragas
LangChain

Job description

UNAVAILABLE is seeking an AI Evaluation Scientist to design and execute evaluation programs for predictive and generative AI systems, ensuring accuracy, safety, and alignment with mission requirements. This role will establish trust in AI solutions and support continuous improvement across the AI lifecycle.

The position collaborates with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Evaluation Scientist: Trustworthy Metrics & Testing
AI Evaluation Scientist: Trustworthy Metrics & Testing

Steampunk, Inc. • McLean (VA)

On-site
USD 105,000 - 145,000
AI Safety Evaluator for Real-World AI Systems
AI Safety Evaluator for Real-World AI Systems

mpathic • Seattle (WA)

On-site
USD 70,000 - 110,000
AI Evaluation Scientist
AI Evaluation Scientist

UNAVAILABLE • McLean (VA)

On-site
USD 120,000 - 150,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
Safety Evaluation Engineer: Turn AI Risk into Metrics
Safety Evaluation Engineer: Turn AI Risk into Metrics

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Bonus potential
Equity
Benefits
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • New York (NY)

On-site
USD 120,000 - 190,000
AI Safety Evaluation Expert — Remote Contract
AI Safety Evaluation Expert — Remote Contract

Mercor • San Francisco (CA)

Hybrid
GBP 63,000 - 73,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 210,000