Frontier AI Code Engineer — ML Systems & Evaluation

Mercor

New York (NY)

On-site

USD 179,000 - 276,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on realistic ML engineering workflows and model evaluation.

From sprint-based tasks to edge-case identification and performance reviews, you will use frontier AI coding agents to complete complex ML tasks, review model implementations, and compare outputs across frontier models.

Qualifications

  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated machine learning implementations and technical tradeoffs.
  • Experience deploying ML systems to production is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks.
  • Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios.

Skills

ML engineering
Production ML systems
AI coding agents
Model evaluation
ML deployment

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on realistic ML engineering workflows and model evaluation.

From sprint-based tasks to edge-case identification and performance reviews, you will use frontier AI coding agents to complete complex ML tasks, review model implementations, and compare outputs across frontier models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Frontier AI Data Engineer: ETL & Model Evaluation
Frontier AI Data Engineer: ETL & Model Evaluation

Mercor • Philadelphia

On-site
USD 179,000 - 276,000
Frontier AI Data Engineer — Model Evaluation & ETL
Frontier AI Data Engineer — Model Evaluation & ETL

Mercor • San Francisco (CA)

On-site
USD 207,000 - 234,000
Frontier ML Engineer: AI Coding Agent Evaluator
Frontier ML Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

Hybrid
AI-Driven DevOps Evaluator for Frontier Coding Agents
AI-Driven DevOps Evaluator for Frontier Coding Agents

Mercor • New York (NY)

On-site
USD 441,000 - 661,000
Frontier AI Infrastructure Engineer (Contract)
Frontier AI Infrastructure Engineer (Contract)

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000
Frontier AI ML Engineer — Evaluation & Production Systems
Frontier AI ML Engineer — Evaluation & Production Systems

Obsidian • New York (NY)

Remote
USD 4,000 - 7,000
Frontier AI ML Engineer — Real-World Model Evaluator
Frontier AI ML Engineer — Real-World Model Evaluator

Great Value Hiring • United States

On-site
ML Engineer - AI Coding Expert - AI Trainer
ML Engineer - AI Coding Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Frontier ML Engineer: Benchmark AI Coding Models
Frontier ML Engineer: Benchmark AI Coding Models

Obsidian • New York (NY)

On-site
USD 207,000 - 303,000