Machine Learning Engineer — Model Evaluation & Experimentation

Dorado

United States

Remote

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor, in partnership with Cincinnatus LLC, is hiring experienced machine learning researchers to author and execute complex, multi-step ML tasks for frontier models. You will design, implement, train, and analyze experiments, focusing on evaluation and improvements in reinforcement learning concepts.

You’ll work remotely within the United States, committing to approximately 35 hours per week, and collaborate with researchers to ensure rigorous and fair task benchmarks.

Qualifications

  • Master's or PhD in ML, CS or related field, or equivalent research experience.
  • 1+ years in a research or research-engineering role.
  • Hands-on ML model training, evaluation, and end-to-end experimentation experience.
  • Strong familiarity with large language models and evaluation techniques.
  • Proficient in Python and Git; comfortable with notebooks and scripting.
  • Basic RL concepts incl. reward functions and policy training.
  • Prior AI training, model evaluation, or benchmark task authoring is a plus.
  • Attention to detail, creativity in task design, and ability to work independently.

Responsibilities

  • Design multi-step ML tasks from real research ideas, including implementation and evaluation.
  • Run experiments: implement changes, train models, analyze results to define correct solutions.
  • Explore RL ideas around reward functions and training behavior.
  • Evaluate frontier models and document where they fall short.
  • Collaborate with researchers to ensure tasks are rigorous, fair, and consistent.

Skills

ML research
Experiment design
RL basics
LLMs evaluation
Python
Git
Communication
Independent work

Education

MSc or PhD in ML/CS or equivalent
1+ years in a research role

Tools

PyTorch
TensorFlow
Jupyter

Job description

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short.

Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities
  • Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks.

  • Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like.

  • Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior.

  • Evaluate models: See how frontier models handle your tasks, and note where and why they fall short.

  • Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair.

3. Core Qualifications
  • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.

  • 1+ years of experience in a research or research-engineering role.

  • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.

  • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.

  • Working proficiency in Python and Git, with comfort in both scripting and notebook environments.

  • Basic understanding of reinforcement learning (reward functions, policy training) is preferred.

  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.

  • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client\'s internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Expert - AI/ML
Data Science Expert - AI/ML

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Dorado • United States

Remote
USD 90,000 - 120,000
STEM Researcher — Computational Fields
STEM Researcher — Computational Fields

Dorado • United States

Remote
USD 105,000 - 150,000
LLM Red Team Specialist — Failure Modes & Edge Cases
LLM Red Team Specialist — Failure Modes & Edge Cases

Dorado • United States

Remote
USD 120,000 - 180,000
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →

Dorado • California (MO), Northern (KY)

Hybrid
USD 180,000 - 240,000
QA/Test Engineer
QA/Test Engineer

Dorado • United States

Remote
USD 90,000 - 130,000
Software Domain Expert
Software Domain Expert

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
GenAI Model Evaluation Engineer — Remote, 35h/wk
GenAI Model Evaluation Engineer — Remote, 35h/wk

Dorado • United States

Remote
USD 120,000 - 180,000
ML Systems Engineer - AI Trainer
ML Systems Engineer - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 150,000
Finance SME - AI Evaluation Expert
Finance SME - AI Evaluation Expert

Mercor • New York (NY)

On-site
USD 150,000 - 210,000