Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Weekday AI is seeking experienced Machine Learning Engineers and Researchers to design and evaluate frontier AI benchmarks. This fully remote, full-time role involves creating multi-step tasks, running training pipelines, and analyzing model behavior.

You will collaborate with researchers, implement in Python, and contribute to rigorous evaluation methods. The position requires 1+ year ML experience, proficiency in Python and Git, and about 35 hours per week.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently.
  • Strong written communication skills for documenting experimental methodologies and technical findings.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Responsibilities

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency.

Skills

Python
Git
Machine learning
Experimentation
RL concepts
LLMs
Documentation

Education

Master's degree or PhD

Job description

Weekday AI is seeking experienced Machine Learning Engineers and Researchers to design and evaluate frontier AI benchmarks. This fully remote, full-time role involves creating multi-step tasks, running training pipelines, and analyzing model behavior.

You will collaborate with researchers, implement in Python, and contribute to rigorous evaluation methods. The position requires 1+ year ML experience, proficiency in Python and Git, and about 35 hours per week.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote ML Benchmark Engineer & Evaluation Specialist
Remote ML Benchmark Engineer & Evaluation Specialist

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote Data Science & AI Evaluation Benchmark Designer
Remote Data Science & AI Evaluation Benchmark Designer

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Benchmark Engineer (Python) — Remote
AI Benchmark Engineer (Python) — Remote

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote ML Engineer: Build Benchmarks & Production Pipelines
Remote ML Engineer: Build Benchmarks & Production Pipelines

raydar • Northern (KY)

Hybrid
USD 170,000 - 270,000
Competitive equity
Remote Data Science & AI Benchmark Analyst
Remote Data Science & AI Benchmark Analyst

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking

YO HR Consultancy • United States

Remote
USD 110,208 - 165,312
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday AI • United States

Remote
USD 83,000 - 124,000
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Frontier AI Benchmark Researcher — Remote
Frontier AI Benchmark Researcher — Remote

Synthires • United States

Remote
USD 193,000 - 207,000
Fully remote opportunity
Flexible workload (5–40 hours/week)
Fully asynchronous work environment