AI Evaluation Scientist: Trustworthy Metrics & Testing

Steampunk, Inc.

McLean (VA)

On-site

USD 105,000 - 145,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Steampunk, Inc. is seeking an AI Evaluation Scientist to design and execute evaluation programs, ensuring predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements.

You’ll work with engineers, data scientists, and governance analysts to develop evaluation metrics and reports. The role involves building test harnesses, conducting structured error analyses, and contributing to human-in-the-loop workflows to support responsible deployment.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field.
  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
  • Proficiency in AI evaluation frameworks such as Ragas or DeepEval.
  • Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.
  • Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.

Responsibilities

  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
  • Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences.
  • You will contribute to the growth of our AI & Data Exploitation Practice!

Skills

Analytical thinking
Excellent communication
Agile mindset
Public trust eligibility

Education

Bachelor’s or Master’s degree in CS/Stats/ML/AI

Tools

Python
PyTorch
Hugging Face
scikit-learn
LangChain
OpenTelemetry

Job description

Steampunk, Inc. is seeking an AI Evaluation Scientist to design and execute evaluation programs, ensuring predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements.

You’ll work with engineers, data scientists, and governance analysts to develop evaluation metrics and reports. The role involves building test harnesses, conducting structured error analyses, and contributing to human-in-the-loop workflows to support responsible deployment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 105,000 - 145,000
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Evaluation & Reliability Architect
AI Evaluation & Reliability Architect

Landing Point • Village of Pelham (NY)

On-site
USD 152,000 - 179,000
Research Engineer: AI Evaluation & Metrics
Research Engineer: AI Evaluation & Metrics

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
AI Evaluation Scientist — Real-World ML Research
AI Evaluation Scientist — Real-World ML Research

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Health benefits
+2
Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
AI QA & Evaluation Engineer - Scale & Metrics
AI QA & Evaluation Engineer - Scale & Metrics

Appnovation Technologies • New York (NY)

On-site
USD 90,000 - 150,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000