Senior AI/ML Evaluation & Benchmark Engineer

OpenTeams

Colorado

Hybrid

USD 145,000 - 250,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) Match
Unlimited PTO
Fully Remote Setup
Education reimbursement
Disability & life insurance
HSA & FSA

Job summary

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking and evaluation capabilities at the core of our AI platform. This role focuses on what models get wrong and requires hands-on engineering on an open-source toolchain.

The candidate will design evaluation methodologies, document limitations, and surface failure modes to inform decisions about field readiness. Travel up to 15% may be required for government facilities and company locations.

Qualifications

  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance.
  • 6+ years of software or ML engineering experience, with 3+ years evaluating or deploying ML models.
  • Strong Python proficiency in a machine learning or data science context.
  • Hands-on experience with PyTorch and Hugging Face ecosystems.
  • Experience developing or using model evaluation harnesses, benchmark suites, or test & evaluation frameworks.
  • Experience designing evaluation metrics with statistical rigor.
  • Experience building auditable evaluation pipelines with documented data provenance.
  • Experience evaluating large language models or agentic workflows using task-based or judgment-based scoring.
  • Strong written communication for technical and non-technical stakeholders.
  • Bachelor’s degree in CS/Math/Engineering or equivalent practical experience.

Responsibilities

  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows.
  • Develop evaluation methodologies that combine automated metrics with structured human judgment.
  • Curate and recommend candidate benchmarks based on mission needs and document data provenance.
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including limitations.
  • Define and contribute to common standards for benchmark expression, ingestion, and reporting.
  • Support partner organizations and vendors as they integrate with shared evaluation standards.
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement.
  • Participate in structured feedback sessions with mission end users and incorporate findings into the platform.
  • Develop reference notebooks and workflows that enable data-science-capable analysts to run and interpret evaluations.
  • Document technical approaches, evaluation results, and key decisions for Government stakeholders.

Skills

Python
PyTorch
Hugging Face
Model evaluation
Data provenance
Communication
Security clearance
Software engineering

Education

Bachelor’s degree or equivalent

Tools

Jupyter
Benchmarking toolchains

Job description

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking and evaluation capabilities at the core of our AI platform. This role focuses on what models get wrong and requires hands-on engineering on an open-source toolchain.

The candidate will design evaluation methodologies, document limitations, and surface failure modes to inform decisions about field readiness. Travel up to 15% may be required for government facilities and company locations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI/ML Benchmarking & Evaluation Engineer
Senior AI/ML Benchmarking & Evaluation Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Senior AI/ML Test and Evaluation Engineer
Senior AI/ML Test and Evaluation Engineer

openteams • Washington

Hybrid
USD 145,000 - 250,000
Senior AI/ML Test and Evaluation Engineer
Senior AI/ML Test and Evaluation Engineer

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote
Senior AI/ML Test and Evaluation Engineer Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote

OpenTeams • Colorado

Hybrid
USD 145,000 - 250,000
401(k) Match
Unlimited PTO
Fully Remote Setup
+3
Senior AI Test & Evaluation Engineer
Senior AI Test & Evaluation Engineer

Motion • Birmingham (AL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior AI Systems Architect: Agentic & Open-Source ML
Senior AI Systems Architect: Agentic & Open-Source ML

Advantest • Oregon (WI)

On-site
USD 180,000 - 280,000
Senior AI Tooling & Testing Engineer
Senior AI Tooling & Testing Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 147,000 - 238,000
Senior ML Engineer: AI Evaluation & Reproducible Infra
Senior ML Engineer: AI Evaluation & Reproducible Infra

Society for Human Resource Management (SHRM) • Alexandria (VA)

Hybrid
USD 100,000 - 130,000
Health benefits
Dental benefits
Vision benefits
+5
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1