Remote ML Benchmark Engineer & Evaluation Specialist

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday 1 is seeking experienced ML Engineers and Researchers for a fully remote, independent contractor role focused on building evaluation benchmarks for frontier AI models. You will design multi-step benchmarks, implement experiments, and analyze model behavior.

The position is about 35 hours per week, paid hourly at $60-$90, with weekly payments. It requires strong Python/Git skills, ML research experience, and familiarity with LLMs and reinforcement learning.

Qualifications

  • Master's degree, PhD, or equivalent in a quantitative STEM discipline.
  • Minimum 1 year of professional experience in ML research or related field.
  • Hands‑on experience designing, training, and evaluating ML models through full experimental workflows.
  • Experience with ML experiments: setup, tuning, execution, validation, analysis.
  • Strong understanding of modern LLMs and evaluation methodologies.
  • Proficiency in Python and Git; experience in script-based and notebook environments.
  • Familiarity with reinforcement learning concepts is preferred.
  • Experience in AI evaluation, benchmark development, or research experimentation is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and ability to work independently.

Responsibilities

  • Design realistic ML benchmark tasks based on research workflows.
  • Translate research concepts into structured, reproducible evaluation tasks.
  • Implement ML solutions in Python, run experiments, and provide reference implementations.
  • Develop benchmark tasks involving reinforcement learning concepts where applicable.
  • Evaluate AI solutions by identifying errors, flaws, and unsupported conclusions.
  • Collaborate with researchers to improve benchmark quality and evaluation consistency.

Skills

LLMs
Reinforcement learning
Experimentation
Analytical thinking

Education

Master's degree, PhD, or equivalent in ML/CS/AI/Data Science

Tools

Python
Git

Job description

Weekday 1 is seeking experienced ML Engineers and Researchers for a fully remote, independent contractor role focused on building evaluation benchmarks for frontier AI models. You will design multi-step benchmarks, implement experiments, and analyze model behavior.

The position is about 35 hours per week, paid hourly at $60-$90, with weekly payments. It requires strong Python/Git skills, ML research experience, and familiarity with LLMs and reinforcement learning.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Benchmark Engineer (Python) — Remote
AI Benchmark Engineer (Python) — Remote

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote ML & NLP Expert for Evaluation & R&D
Remote ML & NLP Expert for Evaluation & R&D

Weekday 1 • United States

Remote
USD 110,000 - 152,000
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking

YO HR Consultancy • United States

Remote
USD 110,208 - 165,312
AI Benchmark Architect (Remote) | Python
AI Benchmark Architect (Remote) | Python

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Remote Data Science & AI Evaluation Benchmark Designer
Remote Data Science & AI Evaluation Benchmark Designer

Weekday AI • United States

Remote
USD 83,000 - 124,000
QA Engineer - AI Benchmarking (Remote, 35h/wk)
QA Engineer - AI Benchmarking (Remote, 35h/wk)

Weekday AI • United States

Remote
USD 83,000 - 124,000