ML Engineer - Coding Agent Expert

Obsidian

New York (NY)

Hybrid

USD 275,520 - 826,560

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking contributors for a Frontier Code Agents project focused on evaluating and improving AI coding models. Candidates will complete and assess complex machine learning and AI engineering tasks.

The role requires at least 2 years of experience in machine learning engineering, knowledge of AI coding agents, and technical evaluation skills. Compensation is performance-based at $400 per accepted task.

Qualifications

  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems and applications.
  • Ability to evaluate model-generated implementations.

Responsibilities

  • Use frontier AI coding agents to complete complex ML tasks.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple models to assess strengths.

Skills

Machine learning engineering
AI coding agents
Model evaluation
Production ML systems

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic machine learning engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks.
  • Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios.
Time Commitment

Sprint based project that runs in 12-24 hour stretches based on client requirement.

Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated machine learning implementations and technical tradeoffs.
  • Experience deploying ML systems to production is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Mercor • New York (NY)

On-site
USD 179,000 - 276,000
ML Engineer - AI Coding Expert - AI Trainer
ML Engineer - AI Coding Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Data Engineer - AI Model Evaluator - AI Trainer
Data Engineer - AI Model Evaluator - AI Trainer

Mercor • Philadelphia

On-site
USD 179,000 - 276,000
Fraud Engineer - AI Specialist
Fraud Engineer - AI Specialist

Obsidian • San Francisco (CA)

On-site
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • San Francisco (CA)

On-site
USD 207,000 - 551,000
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Mercor • San Francisco (CA)

On-site
USD 207,000 - 234,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Data Engineer - Fully Remote | Upto $80/hr
Data Engineer - Fully Remote | Upto $80/hr

Obsidian • San Francisco (CA)

Remote
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000