ML Research Engineer: End-to-End AI Evaluation & Experiments

24-Mag Llc

New York (NY)

Remote

USD 100,000 - 155,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

24-MAG LLC is offering a fully remote, full-time opportunity for machine learning engineers and research practitioners to design, implement, and evaluate end-to-end ML benchmarks.

You will transform real ML ideas into multi-step tasks, run experiments in Python notebooks, analyze training behavior, and assess model-generated solutions for correctness. Collaboration with researchers and task authors is expected, with ~35 hours/week and competitive hourly rates.

Qualifications

  • Master's degree or PhD in machine learning, computer science, AI, engineering, mathematics, or related STEM field.
  • Equivalent practical experience in a research-intensive ML role may be considered.
  • Publications, open-source contributions, technical reports, or substantial research projects are valuable.

Responsibilities

  • Design multi-step evaluation tasks from practical ML research ideas.
  • Implement reference solutions using Python, scripts, and notebooks.
  • Configure and run training experiments from setup to evaluation.
  • Validate code, dependencies, datasets, outputs, and results.
  • Document workflows to enable reproducibility.

Skills

Python
Git
Experiment design
Reproducible research
Notebook environments
Analytical reasoning
Communication

Education

Master's degree or PhD in ML/CS/AI
Equivalent practical experience in ML research
Publications/open-source contributions

Tools

Python
Jupyter

Job description

24-MAG LLC is offering a fully remote, full-time opportunity for machine learning engineers and research practitioners to design, implement, and evaluate end-to-end ML benchmarks.

You will transform real ML ideas into multi-step tasks, run experiments in Python notebooks, analyze training behavior, and assess model-generated solutions for correctness. Collaboration with researchers and task authors is expected, with ~35 hours/week and competitive hourly rates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote CS Research Expert - Benchmark & ML Systems
Remote CS Research Expert - Benchmark & ML Systems

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 103,000
Remote work
Flexible schedule
Competitive hourly rate
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
Senior ML Engineer - Safety & AI Evaluation (Remote)
Senior ML Engineer - Safety & AI Evaluation (Remote)

10a Labs • Chicago (IL)

On-site
USD 130,000 - 200,000
Performance-based annual bonus
Comprehensive health, dental, and vision coverage
Support for conferences and continuing education
+1
Machine Learning Expert - Fully Remote | Upto $90/hr
Machine Learning Expert - Fully Remote | Upto $90/hr

Obsidian • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
ML Research Engineer: Fine-Tuning & Evaluation Systems
ML Research Engineer: Fine-Tuning & Evaluation Systems

HonestAI • San Francisco (CA)

On-site
USD 140,000 - 190,000
Remote Data Scientist & Quant Analyst for AI Evaluation
Remote Data Scientist & Quant Analyst for AI Evaluation

24-Mag Llc • New York (NY)

Remote
Remote Python Engineer for AI Benchmarking
Remote Python Engineer for AI Benchmarking

24-Mag Llc • New York (NY)

Remote
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking
Remote ML Engineer — Part-Time (20h/wk) for AI Benchmarking

YO HR Consultancy • United States

Remote
USD 125,000 - 150,000
ML Safety & Benchmarking Research Engineer
ML Safety & Benchmarking Research Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 181,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Senior ML Engineer - Safety & AI Evaluation (Remote)
Senior ML Engineer - Safety & AI Evaluation (Remote)

10a Labs • New York (NY)

Remote
USD 130,000 - 200,000
Comprehensive health, dental, and vision coverage
Performance-based annual bonus
Support for conferences and continuing education