AI Evaluation Scientist: Trusted AI Metrics & Tests

Steampunk, Inc.

McLean (VA)

On-site

USD 105,000 - 145,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Steampunk seeks an AI Evaluation Scientist to design and execute evaluation processes ensuring AI systems are accurate, reliable, safe, and aligned with mission requirements.

You will collaborate with AI Developers, LLMOps Engineers, and Data Scientists to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment across the AI lifecycle.

Qualifications

  • Bachelor’s degree and 5+ years of experience.
  • Master’s degree and 3+ years of experience.
  • 2+ years evaluating ML models, NLP systems, or generative AI models (LLMs preferred).
  • Proficiency in Python and relevant libraries (PyTorch, Hugging Face, scikit-learn, LangChain).
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design.

Responsibilities

  • Design evaluation processes to ensure accuracy, reliability, safety, and alignment with mission requirements.
  • Develop benchmarks, test harnesses, and pipelines to monitor performance drift.
  • Perform structured error analysis and audits of LLMs, RAG, and predictive models.

Skills

Python
PyTorch
Hugging Face
scikit-learn
LangChain

Education

Bachelor's degree in CS/Stats/ML/etc.
Master's degree in CS/Stats/ML/etc.

Tools

OpenTelemetry
Ragas/DeepEval

Job description

Steampunk seeks an AI Evaluation Scientist to design and execute evaluation processes ensuring AI systems are accurate, reliable, safe, and aligned with mission requirements.

You will collaborate with AI Developers, LLMOps Engineers, and Data Scientists to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment across the AI lifecycle.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 105,000 - 145,000
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Senior AI Evaluation Engineer: Pipelines & Metrics
Senior AI Evaluation Engineer: Pipelines & Metrics

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
QA Engineer – AI/LLM Testing & DevSecOps
QA Engineer – AI/LLM Testing & DevSecOps

Steampunk, Inc. • McLean (VA)

On-site
USD 90,000 - 130,000
AI QA & Evaluation Engineer - Scale & Metrics
AI QA & Evaluation Engineer - Scale & Metrics

Appnovation Technologies • New York (NY)

On-site
USD 90,000 - 150,000
AI Model Evaluation Lead: Metrics, Bias & Fairness
AI Model Evaluation Lead: Metrics, Bias & Fairness

MERIT Beauty • New York (NY)

On-site
AI Quality Engineer: Evaluation Pipelines & LLM Testing
AI Quality Engineer: Evaluation Pipelines & LLM Testing

Dealstitch LLC. • United States

On-site
USD 160,000 - 220,000
Senior AI Tech Lead for Enterprise ML & AI Systems
Senior AI Tech Lead for Enterprise ML & AI Systems

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 180,000