AI Evaluation Scientist

Steampunk

McLean (VA)

On-site

USD 105,000 - 145,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Steampunk is seeking an AI Evaluation Scientist in McLean, Virginia, to design evaluation processes for AI systems ensuring accuracy and reliability. Responsibilities include implementing evaluation frameworks and performing error analysis. Ideal candidates will have a Bachelor’s or Master’s degree in a relevant field, with 2+ years of experience in AI model evaluation and proficiency in Python. The expected salary range is $105,000 to $145,000 annually, backed by a comprehensive benefits package.

Qualifications

  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity with dataset creation and experimental design for AI systems.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.

Responsibilities

  • Implement evaluation frameworks for AI models to ensure reliability and accuracy.
  • Build and maintain automated evaluation scripts and pipelines.
  • Perform structured error analysis and behavioral audits of models.

Skills

Experience evaluating machine learning models
Proficiency in Python and relevant libraries
Strong analytical skills
Excellent communication skills
Familiarity with evaluation metrics
Understanding of LLM behavior

Education

Bachelor’s or Master’s degree in a relevant field

Tools

Python
PyTorch
Hugging Face
scikit-learn
LangChain

Job description

Overview

We are looking for anAI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.

Contributions
  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
  • Contribute to the design of human‑in‑the‑loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
  • Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
  • Document evaluation processes, criteria, and results for both technical and non‑technical audiences.
  • You will contribute to the growth of our AI & Data Exploitation Practice!
Qualifications
  • Ability to hold a position of public trust with the U.S. government.
  • Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human‑Computer Interaction, Data Science, or a related field.
  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit‑learn, LangChain.
  • Proficiency in AI evaluation frameworks such as Ragas.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.
  • Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
  • Exposure to responsible AI concepts and governance‑aligned evaluation criteria (e.g., fairness, transparency, reliability).
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non‑technical stakeholders.
  • Experience working in agile or iterative development environments is a plus.
  • Familiarity with OWASP LLM Top 10 Risks.
  • NIH experience.
  • Local to Washington, DC metro area preferred.
  • Relevant certifications (helpful but not required):
    • NIST AI RMF (AISIC)
    • INFORMS CAP
    • AWS/Azure/Google ML Certifications.
Compensation & Benefits

Steampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000. The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk’s total compensation package for employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLMOps Engineer
LLMOps Engineer

Steampunk • McLean (VA)

On-site
USD 115,000 - 145,000
Senior AI Developer
Senior AI Developer

Steampunk • McLean (VA)

On-site
USD 140,000 - 180,000
Senior LLMOps Engineer
Senior LLMOps Engineer

Steampunk, Inc. • McLean (VA)

On-site
USD 145,000 - 185,000
AI Developer
AI Developer

Steampunk • McLean (VA)

On-site
USD 140,000 - 190,000
Employee ownership
Comprehensive benefits package
Professional development opportunities
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Data Scientist (Generative AI)
Data Scientist (Generative AI)

Steampunk, Inc. • McLean (VA)

On-site
USD 125,000 - 160,000
MLOps Engineer
MLOps Engineer

Steampunk, Inc. • McLean (VA)

On-site
USD 115,000 - 150,000
Prompt Engineer
Prompt Engineer

Steampunk • McLean (VA)

On-site
USD 115,000 - 140,000
Data Scientist
Data Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 160,000
AI/ML Engineer, Senior
AI/ML Engineer, Senior

Booz Allen Hamilton • Dayton (OH)

On-site
USD 99,000 - 225,000
Health insurance
Life and disability insurance
Retirement benefits
+1