Remote Part-Time ML Engineer — Evaluation & Experiments

OpenTrain AI

United States

On-site

USD 82,656 - 123,984

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenTrain AI is seeking an experienced machine learning practitioner to act as a ground-truth expert for frontier model evaluation. You will turn research ideas into reproducible evaluation tasks, run training experiments, and analyze results to show correct solutions and where models fall short.

This is a contract, part-time position for US-based contributors, expected at 20+ hours per week.

Qualifications

  • MSc/PhD in ML, CS or related STEM, or equivalent experience.
  • At least 1 year in research or research-engineering role.
  • Hands-on experience training and evaluating ML models end-to-end.
  • Strong familiarity with large language models, their capabilities, limitations, and evaluation techniques.
  • Proficient in Python and Git; comfortable in scripting and notebook environments.

Responsibilities

  • Design, implement, and run end-to-end evaluation and experimentation workflows that benchmark language models and training approaches.
  • Turn vague ML research ideas into well-defined, multi-step evaluation tasks and benchmarks.
  • Implement changes, run training experiments, and analyze outcomes to demonstrate correct behavior.
  • Build tasks and analyses around reinforcement learning concepts such as reward functions and training behavior.
  • Evaluate frontier models on your tasks, document failure modes, and explain where and why they fall short.
  • Collaborate with researchers and other experts to keep tasks consistent, rigorous, and reproducible.

Skills

Research problem solving
Written communication
Independent work
Experiment design

Education

MSc or PhD in ML/CS or related STEM (or equivalent experience)
1+ years in research or research-engineering role

Tools

Python
Git

Job description

OpenTrain AI is seeking an experienced machine learning practitioner to act as a ground-truth expert for frontier model evaluation. You will turn research ideas into reproducible evaluation tasks, run training experiments, and analyze results to show correct solutions and where models fall short.

This is a contract, part-time position for US-based contributors, expected at 20+ hours per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer — Model Evaluation & Experimentation
Machine Learning Engineer — Model Evaluation & Experimentation

OpenTrain AI • United States

On-site
Remote ML/AI Expert (PhD) - Model Evaluation
Remote ML/AI Expert (PhD) - Model Evaluation

United States Digital Space LLC • Germany (OH)

On-site
USD 165,000 - 248,000
Flexible hours
ML Research Engineer: End-to-End AI Evaluation & Experiments
ML Research Engineer: End-to-End AI Evaluation & Experiments

24-Mag Llc • New York (NY)

Remote
USD 100,000 - 155,000
Remote ML Model Evaluation & QA Specialist
Remote ML Model Evaluation & QA Specialist

AuraOne • United States

On-site
USD 83,000 - 124,000
Machine Learning Expert - Fully Remote | Upto $90/hr
Machine Learning Expert - Fully Remote | Upto $90/hr

Obsidian • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Remote Member of Technical Staff — Frontier AI (ML Evaluation)
Remote Member of Technical Staff — Frontier AI (ML Evaluation)

Crossing Hurdles • United States

On-site
USD 600,000 - 2,000,000
Remote Part-Time: Personalized AI Response Evaluator
Remote Part-Time: Personalized AI Response Evaluator

OpenTrain AI • United States

On-site
USD 28,000 - 41,000
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
Remote AI Data Scientist — RLHF & Evaluation
Remote AI Data Scientist — RLHF & Evaluation

OpenTrain AI • United States

On-site
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking

YO HR Consultancy • United States

Remote
USD 125,000 - 150,000