Frontier AI ML Engineer — Evaluation & Production Systems

Obsidian

New York (NY)

Remote

USD 4,000 - 7,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking contributors to evaluate frontier AI coding models through structured technical assessments. You will work on realistic ML engineering workflows, model evaluation, and production-ready deployment considerations.

The role emphasizes hands-on use of frontier agents, reviewing model-implemented systems, and applying engineering judgment to complex ML scenarios. Collaboration and rapid iteration are key.

Qualifications

  • 2+ years of professional ML engineering experience.
  • Experience building production ML systems and deployment infra.
  • Experience using AI coding agents.
  • Ability to evaluate model-generated ML implementations.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex ML tasks.
  • Review model-generated implementations including training, inference, MLOps, and LLM apps.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from frontier models and assess strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML scenarios.

Skills

ML engineering
Model evaluation
Debugging performance
MLOps
AI tooling

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Obsidian is seeking contributors to evaluate frontier AI coding models through structured technical assessments. You will work on realistic ML engineering workflows, model evaluation, and production-ready deployment considerations.

The role emphasizes hands-on use of frontier agents, reviewing model-implemented systems, and applying engineering judgment to complex ML scenarios. Collaboration and rapid iteration are key.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier ML Engineer: Benchmark AI Coding Models
Frontier ML Engineer: Benchmark AI Coding Models

Obsidian • New York (NY)

On-site
USD 207,000 - 303,000
Frontier ML Engineer: AI Coding Agent Evaluator
Frontier ML Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

Hybrid
AI Code-Agent Evaluator: Frontier DevOps Engineer
AI Code-Agent Evaluator: Frontier DevOps Engineer

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
AI-Driven Data Engineer - Frontier Pipeline Evaluator
AI-Driven Data Engineer - Frontier Pipeline Evaluator

Obsidian • New York (NY)

Remote
USD 193,000 - 248,000
AI Systems Engineer: Frontier Model Evaluator
AI Systems Engineer: Frontier Model Evaluator

Obsidian • New York (NY)

On-site
USD 200,000
Frontier AI ML Engineer — Real-World Model Evaluator
Frontier AI ML Engineer — Real-World Model Evaluator

Great Value Hiring • United States

On-site
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Frontier Cloud Engineer: AI Coding Agents
Frontier Cloud Engineer: AI Coding Agents

Obsidian • New York (NY)

Remote
USD 551,000 - 1,102,000
Frontier AI Data Engineer: Model Evaluation & Pipelines
Frontier AI Data Engineer: Model Evaluation & Pipelines

Obsidian • Philadelphia

On-site
USD 455,000 - 647,000
Frontier AI Data Engineer: ETL & Model Evaluation
Frontier AI Data Engineer: ETL & Model Evaluation

Mercor • Philadelphia

On-site
USD 179,000 - 276,000