Remote Senior AI/ML Benchmarking & Evaluation Engineer

OpenTeams

Northern (KY)

Hybrid

USD 145,000 - 250,000

Full time

8 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

OpenTeams seeks a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking and evaluation capabilities at the core of our AI platform. You will design automated metrics and human judgments to compare models and agentic workflows, surface failure modes, and ensure evaluations stay valid as models evolve.

This hands-on role uses an open‑source toolchain, reports to senior stakeholders, and helps decide which capabilities are ready to field, emphasizing robust, transparently

Qualifications

  • Experience designing benchmarks for ML/agentic systems.
  • Proficiency with Python for API frameworks and evaluation harnesses.
  • Familiarity with ML/vision/language models and evaluation.

Responsibilities

  • Design and implement a platform for test & evaluation of AI models and agentic systems.
  • Develop evaluation methodologies combining human and AI expert judging and multiple modalities.
  • Design and enhance user interfaces for workflow expression by customers.
  • Collaborate with SMEs, model/agent developers, and evaluation designers.
  • Architect and implement robust experiment provenance and result tracking.
  • Ensure performance across integrated tools and scaling systems.
  • Provide mentorship, code reviews, and documentation for knowledge transfer.

Skills

ML platforms
Python API
Computer vision / language modeling
PyTorch / HuggingFace
Cross-functional collaboration
Technical leadership

Tools

PyTorch
Hugging Face

Job description

OpenTeams seeks a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking and evaluation capabilities at the core of our AI platform. You will design automated metrics and human judgments to compare models and agentic workflows, surface failure modes, and ensure evaluations stay valid as models evolve.

This hands-on role uses an open‑source toolchain, reports to senior stakeholders, and helps decide which capabilities are ready to field, emphasizing robust, transparently

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Evaluation Scientist — Benchmarks & Systems
Senior AI Evaluation Scientist — Benchmarks & Systems

Oracle • Santa Clara (CA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off and holidays
+1
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior AI/ML Test and Evaluation Engineer United States - Remote OR Hybrid
Senior AI/ML Test and Evaluation Engineer United States - Remote OR Hybrid

OpenTeams • Northern (KY)

Hybrid
USD 145,000 - 250,000
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Senior Applied Scientist: AI Evaluation & Benchmarks
Senior Applied Scientist: AI Evaluation & Benchmarks

Oracle • United States

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Short/Long-term disability
Life insurance & AD&D
+2
Remote AI Model Evaluator & Benchmark Scientist
Remote AI Model Evaluator & Benchmark Scientist

OpenTrain AI, Inc. • United States

Remote
USD 34,000 - 55,000
Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
Senior AI Benchmarking & Systems Architect
Senior AI Benchmarking & Systems Architect

Aionia Group • San Francisco (CA)

On-site
USD 130,000 - 220,000
Equity
On-site
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1