ML Engineer - AI Coding Expert

Mercor

Toronto

On-site

CAD 179,000 - 289,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor collaborates with a leading AI research lab on frontier code agents to assess and improve frontier AI coding models through structured evaluations. Contributors help advance realistic ML engineering workflows and model evaluation, focusing on production-ready deployment and robust inference systems.

The role emphasizes hands-on review of model-generated implementations, bug-hunting, and comparing outputs across frontier models, requiring strong engineering judgment for real-world ML

Qualifications

  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems, deployment infra, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated ML implementations and technical tradeoffs.
  • Experience deploying ML systems to production is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex ML and AI engineering tasks.
  • Review model-generated implementations involving training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple frontier models and assess strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios.

Skills

ML engineering
Production ML systems
LLM applications
AI coding agents
Model deployment infrastructure
Evaluation of model outputs
Problem solving

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic machine learning engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks.
  • Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2-3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated machine learning implementations and technical tradeoffs.
  • Experience deploying ML systems to production is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Obsidian • Toronto

On-site
CAD 767,000 - 863,000
Machine Learning Engineer - Fully Remote | Upto $85/hr
Machine Learning Engineer - Fully Remote | Upto $85/hr

Obsidian • Toronto

Remote
CAD 183,000 - 276,000
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • Toronto

On-site
CAD 634,000 - 903,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Toronto

On-site
CAD 18,000 - 27,000
Frontier AI ML Engineer - Evaluation & Production Tasks
Frontier AI ML Engineer - Evaluation & Production Tasks

Obsidian • Toronto

Remote
CAD 183,000 - 276,000
Remote ML Engineer: Frontier AI & Production Systems
Remote ML Engineer: Frontier AI & Production Systems

Mercor • Montreal (administrative region)

On-site
Freelance Software Engineer - AI Coding Agent Evaluation Mindrift · Remote Pay not listed →
Freelance Software Engineer - AI Coding Agent Evaluation Mindrift · Remote Pay not listed →

Dorado • Canada

Hybrid
CAD 34,000 - 48,000
Software Engineering Evaluation Specialist
Software Engineering Evaluation Specialist

Socket.dev • Quebec

On-site
CAD 48,000 - 67,000
Data Engineering Specialist - Fully Remote | Upto $80/hr
Data Engineering Specialist - Fully Remote | Upto $80/hr

NEPSE Trading • Canada

Hybrid
CAD 94,000 - 127,000
Forward Deployed Engineer
Forward Deployed Engineer

Robots & Pencils LP • Canada

On-site
CAD 176,000 - 244,000