Senior AI/ML Benchmarking & Evaluation Engineer

openteams

Washington

Hybrid

USD 145,000 - 250,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer in a hybrid role based in Washington, DC or CO locations. You will build and operate evaluation harnesses and automated metrics alongside expert judgment to stress-test models and agentic workflows.

You will document limitations, surface critical failure modes, and produce defensible reports for senior stakeholders. This hands-on engineering role uses open-source toolchains and may require travel up to 15% to government facilities

Qualifications

  • US citizenship and eligibility to obtain and maintain a US security clearance
  • 6+ years of software engineering or ML engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in prod or applied research
  • Strong Python proficiency in ML/data science context
  • Hands-on experience with ML frameworks and tooling such as PyTorch and Hugging Face ecosystem
  • Experience developing or using model evaluation harnesses, benchmark suites, or test/evaluation frameworks
  • Experience designing evaluation metrics with statistical rigor
  • Experience building repeatable, auditable evaluation pipelines with data provenance
  • Experience evaluating LLMs or agentic workflows using task-, metric-, or judgment-based scoring
  • Strong written communication for explaining evaluation methodologies and results
  • Ability to translate mission needs into practical evaluation approaches

Responsibilities

  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows
  • Develop evaluation methodologies combining automated metrics with structured human expert judgment
  • Curate and recommend candidate benchmarks based on mission needs and document provenance of ground-truth data
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows and document limitations
  • Define and contribute to standards for benchmark expression, ingestion, and reporting
  • Support partner organizations and vendors as they integrate with shared evaluation standards
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement
  • Participate in feedback sessions with mission end users and incorporate findings
  • Develop reference notebooks and example workflows for data-science-capable analysts
  • Document technical approaches, evaluation results, and key decisions for government stakeholders and teams

Skills

Python
PyTorch
Hugging Face
ML Evaluation
Benchmarking
Software Engineering
Security Clearance
US Citizenship

Education

Bachelor's degree in computer science, mathematics, engineering, or related field

Tools

Evaluation harnesses
Benchmark suites

Job description

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer in a hybrid role based in Washington, DC or CO locations. You will build and operate evaluation harnesses and automated metrics alongside expert judgment to stress-test models and agentic workflows.

You will document limitations, surface critical failure modes, and produce defensible reports for senior stakeholders. This hands-on engineering role uses open-source toolchains and may require travel up to 15% to government facilities

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Evaluation & Benchmark Engineer
Senior AI/ML Evaluation & Benchmark Engineer

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI/ML Test and Evaluation Engineer
Senior AI/ML Test and Evaluation Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Senior AI/ML Test and Evaluation Engineer
Senior AI/ML Test and Evaluation Engineer

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI Test & Evaluation Engineer
Senior AI Test & Evaluation Engineer

Motion • Birmingham (AL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior AI Agent Evaluation Engineer
Senior AI Agent Evaluation Engineer

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 357,000
AI/ML Test Engineer
AI/ML Test Engineer

Modern Technology Solutions, Inc. • Springfield (VA)

On-site
USD 80,000 - 120,000
Hybrid Contractor: Senior Engineering & AI Training Expert
Hybrid Contractor: Senior Engineering & AI Training Expert

OpenTrain AI • California (MO), Northern (KY)

Hybrid
USD 90,000 - 145,000
AI Evaluation Engineer - Reproducible ML Benchmarks Hybrid
AI Evaluation Engineer - Reproducible ML Benchmarks Hybrid

SHRM • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Retirement plan
Bonuses & incentives
+1