Senior AI/ML Test and Evaluation Engineer

openteams

Washington

Hybrid

USD 145,000 - 250,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer in a hybrid role based in Washington, DC or CO locations. You will build and operate evaluation harnesses and automated metrics alongside expert judgment to stress-test models and agentic workflows.

You will document limitations, surface critical failure modes, and produce defensible reports for senior stakeholders. This hands-on engineering role uses open-source toolchains and may require travel up to 15% to government facilities

Qualifications

  • US citizenship and eligibility to obtain and maintain a US security clearance
  • 6+ years of software engineering or ML engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in prod or applied research
  • Strong Python proficiency in ML/data science context
  • Hands-on experience with ML frameworks and tooling such as PyTorch and Hugging Face ecosystem
  • Experience developing or using model evaluation harnesses, benchmark suites, or test/evaluation frameworks
  • Experience designing evaluation metrics with statistical rigor
  • Experience building repeatable, auditable evaluation pipelines with data provenance
  • Experience evaluating LLMs or agentic workflows using task-, metric-, or judgment-based scoring
  • Strong written communication for explaining evaluation methodologies and results
  • Ability to translate mission needs into practical evaluation approaches

Responsibilities

  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows
  • Develop evaluation methodologies combining automated metrics with structured human expert judgment
  • Curate and recommend candidate benchmarks based on mission needs and document provenance of ground-truth data
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows and document limitations
  • Define and contribute to standards for benchmark expression, ingestion, and reporting
  • Support partner organizations and vendors as they integrate with shared evaluation standards
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement
  • Participate in feedback sessions with mission end users and incorporate findings
  • Develop reference notebooks and example workflows for data-science-capable analysts
  • Document technical approaches, evaluation results, and key decisions for government stakeholders and teams

Skills

Python
PyTorch
Hugging Face
ML Evaluation
Benchmarking
Software Engineering
Security Clearance
US Citizenship

Education

Bachelor's degree in computer science, mathematics, engineering, or related field

Tools

Evaluation harnesses
Benchmark suites

Job description

Who We Are

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.


OpenTeams exists to make ownership possible.


Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.


If that sounds like your kind of work, we'd like to meet you.


Senior AI/ML Test and Evaluation Engineer

Location: Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered.


Work Authorization: U.S. citizenship required


Clearance: An active U.S. security clearance is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a clearance.


Salary Range: $145,000-$250,000 USD, dependent on experience level and location


About the Role

We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right.


You build the evaluation harnesses - automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about.


Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that.


This is hands-on engineering on an open-source toolchain.


This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations.


Key Responsibilities


  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows

  • Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment

  • Curate and recommend candidate benchmarks based on mission needs and document the provenance of ground-truth and reference data

  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes

  • Define and contribute to common standards for benchmark expression, ingestion, and reporting

  • Support partner organizations and vendors as they integrate their capabilities with shared evaluation standards

  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement for judgment-based evaluations

  • Participate in structured feedback sessions with mission end users and incorporate findings into the platform and evaluation methodology

  • Develop reference notebooks and example workflows that enable data-science-capable analysts to run, interpret, and extend evaluations

  • Document technical approaches, evaluation results, and key decisions for Government stakeholders and internal teams


Required Skills & Experience


  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance

  • 6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments

  • Strong Python proficiency in a machine learning or data science context

  • Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem

  • Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks

  • Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results

  • Experience building repeatable and auditable evaluation pipelines with documented data provenance

  • Experience evaluating large language models or agentic workflows using task-based, metric-based, or judgment-based scoring

  • Strong written communication skills, including the ability to clearly explain evaluation methodologies, results, limitations, and failure modes to technical and nontechnical stakeholders

  • Ability to work effectively in an evolving environment and translate mission needs into practical evaluation approaches

  • Bachelor's degree in computer science, mathematics, engineering, or a related f

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Test and Evaluation Engineer
Senior AI/ML Test and Evaluation Engineer

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Senior AI/ML Benchmarking & Evaluation Engineer
Senior AI/ML Benchmarking & Evaluation Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Senior AI/ML Evaluation & Benchmark Engineer
Senior AI/ML Evaluation & Benchmark Engineer

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Worky • California (MO)

On-site
USD 180,000 - 230,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
Forward Deployed Engineer - Language Models
Forward Deployed Engineer - Language Models

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including **e