Remote ML Evaluation Architect (Part-Time)

anyone-ai

United States

On-site

USD 83,000 - 124,000

Part time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Anyone AI in the United States is seeking experienced Machine Learning Engineers to review and evaluate ML challenges used in model training. You will analyze experiments, datasets, metrics, and pipelines to determine technical soundness and reproducibility.

The role focuses on ensuring challenges reward good ML reasoning, not just brute-force tuning. Remote, part-time, project-based consulting, with a strong emphasis on written feedback and clear recommendations.

Qualifications

  • 3+ years of hands-on applied ML experience.
  • Strong ability to evaluate ML experiments and datasets.
  • Ability to explain findings with written feedback.

Responsibilities

  • Review ML challenges for design quality and solvability.
  • Evaluate datasets for signals and learnability.
  • Identify shortcuts and artifacts in synthetic data.
  • Assess task calibration and difficulty.
  • Review metrics and thresholds for improvements.
  • Detect data leakage and metric gaming.
  • Verify reproducibility across data → model → eval.
  • Provide actionable recommendations to improve tasks.

Skills

Machine learning
Experiment design
Statistical testing
Data analysis
Written feedback

Tools

Python
TensorFlow
PyTorch
Jupyter

Job description

Anyone AI in the United States is seeking experienced Machine Learning Engineers to review and evaluate ML challenges used in model training. You will analyze experiments, datasets, metrics, and pipelines to determine technical soundness and reproducibility.

The role focuses on ensuring challenges reward good ML reasoning, not just brute-force tuning. Remote, part-time, project-based consulting, with a strong emphasis on written feedback and clear recommendations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Part-Time AI Data Science Domain Expert (Remote)
Part-Time AI Data Science Domain Expert (Remote)

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 138,000 - 276,000
Senior Software Engineer – Open Source & SWE-Bench Evaluation
Senior Software Engineer – Open Source & SWE-Bench Evaluation

anyone-ai • United States

Remote
USD 83,000 - 124,000
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking

YO HR Consultancy • United States

Remote
USD 125,000 - 150,000
Senior ML Engineer - Remote Contractor (Part-Time)
Senior ML Engineer - Remote Contractor (Part-Time)

YO AI Labs • Maryland

Remote
USD 40,000 - 60,000
Remote Freelance ML Task Auditor & AI Evaluation Specialist
Remote Freelance ML Task Auditor & AI Evaluation Specialist

Triwill Group • Northern (KY)

Hybrid
USD 83,000 - 165,000
Remote work
Remote | ML Engineer — $100–$150/hour
Remote | ML Engineer — $100–$150/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 138,000 - 207,000
Remote ML Engineer — Model Training & Evaluation
Remote ML Engineer — Model Training & Evaluation

YO AI Labs • New York (NY)

Remote
USD 83,000 - 165,000
Senior AI Engineer: Remote ML Evaluation & RLHF Training
Senior AI Engineer: Remote ML Evaluation & RLHF Training

Prolific • Seattle (WA)

On-site
USD 80,000 - 100,000
Senior AI Evaluation Engineer (Remote, Part-Time)
Senior AI Evaluation Engineer (Remote, Part-Time)

YO AI Labs • Illinois

Remote
USD 70,000 - 110,000
Senior AI/ML Software Engineer (Remote, Part-time)
Senior AI/ML Software Engineer (Remote, Part-time)

YO AI Labs • Chicago (IL)

Remote
USD 96,000 - 138,000