Frontier ML Engineer: AI Coding Agent Evaluator

Obsidian

New York (NY)

Hybrid

USD 275,520 - 826,560

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking contributors for a Frontier Code Agents project focused on evaluating and improving AI coding models. Candidates will complete and assess complex machine learning and AI engineering tasks.

The role requires at least 2 years of experience in machine learning engineering, knowledge of AI coding agents, and technical evaluation skills. Compensation is performance-based at $400 per accepted task.

Qualifications

  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems and applications.
  • Ability to evaluate model-generated implementations.

Responsibilities

  • Use frontier AI coding agents to complete complex ML tasks.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple models to assess strengths.

Skills

Machine learning engineering
AI coding agents
Model evaluation
Production ML systems

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Obsidian is seeking contributors for a Frontier Code Agents project focused on evaluating and improving AI coding models. Candidates will complete and assess complex machine learning and AI engineering tasks.

The role requires at least 2 years of experience in machine learning engineering, knowledge of AI coding agents, and technical evaluation skills. Compensation is performance-based at $400 per accepted task.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Engineer: AI Coding Agent Evaluator
Backend Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

On-site
Frontier AI Security Engineer (Code Agent Evaluator)
Frontier AI Security Engineer (Code Agent Evaluator)

Obsidian • New York (NY)

Remote
Frontier AI Code Engineer — ML Systems & Evaluation
Frontier AI Code Engineer — ML Systems & Evaluation

Mercor • New York (NY)

On-site
USD 179,000 - 276,000
AI Systems Engineer: Frontier Model Evaluator
AI Systems Engineer: Frontier Model Evaluator

Obsidian • New York (NY)

On-site
USD 200,000
AI Code-Agent Evaluator: Frontier DevOps Engineer
AI Code-Agent Evaluator: Frontier DevOps Engineer

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
Frontier ML Engineer: Benchmark AI Coding Models
Frontier ML Engineer: Benchmark AI Coding Models

Obsidian • New York (NY)

On-site
USD 207,000 - 303,000
Frontier AI Risk Engineer - Fraud & Trust Safety Evaluator
Frontier AI Risk Engineer - Fraud & Trust Safety Evaluator

Obsidian • New York (NY)

Remote
Frontier AI Security Engineer: Code Agent Vetting
Frontier AI Security Engineer: Code Agent Vetting

Obsidian • Chicago (IL)

On-site
Frontier AI ML Engineer — Evaluation & Production Systems
Frontier AI ML Engineer — Evaluation & Production Systems

Obsidian • New York (NY)

Remote
USD 4,000 - 7,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000