AI Evaluation Scientist

UNAVAILABLE

McLean (VA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

UNAVAILABLE is seeking an AI Evaluation Scientist to design and execute evaluation programs for predictive and generative AI systems, ensuring accuracy, safety, and alignment with mission requirements. This role will establish trust in AI solutions and support continuous improvement across the AI lifecycle.

The position collaborates with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support

Qualifications

  • Ability to hold a position of public trust with the U.S. government.
  • Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Data Science, or a related field.
  • 2+ years evaluating ML models, NLP systems, or generative AI models (LLMs preferred).
  • Proficiency in Python and libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Experience in structured error analysis and behavioral audits of AI models.

Responsibilities

  • Implement evaluation frameworks for AI models, including accuracy, robustness, bias, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines to detect drift.
  • Develop benchmark datasets, challenge sets, and scenario-based tests.
  • Perform structured error analysis and audits of LLMs and RAG systems.
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists on experiments and model hardening.
  • Contribute to human-in-the-loop evaluation workflows and governance-aligned reports.
  • Map evaluation outcomes to principles such as fairness, transparency, and safety.
  • Work with AI Governance Analysts to support compliance and risk assessments.
  • Stay current with evaluation tools, metrics, and research in LLM assessment.

Skills

Python
PyTorch
Hugging Face
scikit-learn
LangChain
LLMs
NLP
Evaluation metrics
Experiment design
Data analysis

Education

Bachelor's or Master's in CS/Statistics/ML or related field

Tools

Ragas
LangChain

Job description

Overview

We are looking for anAI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative AI systems areaccurate, reliable, safe, and aligned with mission requirements. This role is essential forestablishingtrust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.

Responsibilities
  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build andmaintainautomated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documentingfindingsand improvement recommendations.
  • Collaborate with AI Developers,LLMOpsEngineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
  • Assistin mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences.
  • You will contribute to the growth of our AI & Data Exploitation Practice!
Qualifications
  • Ability to hold aposition of public trustwith the U.S. government.
  • Bachelor’s orMaster’s degreeinComputer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, ora relatedfield.
  • 2+ yearsof experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity withevaluationmetrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Proficiencyin Python and relevant libraries such asPyTorch, Hugging Face, scikit-learn, LangChain.
  • Proficiency in AI evaluation frameworks such as Ragas.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.
  • Understanding ofLLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.
  • Experience working in agile or iterative development environments is a plus.
  • Familiarity with OWASP LLM Top 10 Risks.
  • NIH experience.
  • Relevant certifications (helpful but not required):
    • NIST AI RMF (AISIC)
    • INFORMS CAP
    • AWS/Azure/Google ML Certifications.
  • Local to Washington, DC metro area preferred.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
LLMOps Engineer
LLMOps Engineer

UNAVAILABLE • McLean (VA)

On-site
USD 150,000 - 210,000
AI Developer
AI Developer

UNAVAILABLE • McLean (VA)

On-site
USD 140,000 - 210,000
Senior AI Developer
Senior AI Developer

UNAVAILABLE • McLean (VA)

On-site
USD 180,000 - 240,000