ML Engineer: Frontier AI Coding & Model Evaluation

Obsidian

Greater London

On-site

GBP 408,000 - 612,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking contributors to work on a Frontier Code Agents project, evaluating and improving frontier AI coding models through structured technical assessments. You will focus on realistic ML engineering workflows, model evaluation, and deploying AI-powered products.

Role requires 2+ years in ML engineering, experience with production ML systems, and familiarity with AI coding agents. Sprint-based work with 12-24 hour stretches as per client needs is expected.

Qualifications

  • 2+ years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated machine learning implementations and technical tradeoffs.
  • Experience deploying ML systems to production is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks.
  • Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios.

Skills

Machine learning engineering
Model evaluation
Production ML systems
LLM applications
AI tooling

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Obsidian is seeking contributors to work on a Frontier Code Agents project, evaluating and improving frontier AI coding models through structured technical assessments. You will focus on realistic ML engineering workflows, model evaluation, and deploying AI-powered products.

Role requires 2+ years in ML engineering, experience with production ML systems, and familiarity with AI coding agents. Sprint-based work with 12-24 hour stretches as per client needs is expected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier ML Engineer — AI Coding Model Evaluator (Contract)
Frontier ML Engineer — AI Coding Model Evaluator (Contract)

Obsidian • Greater London

Remote
GBP 136,000 - 204,000
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Mercor • Greater London

On-site
GBP 184,000 - 255,000
Remote Data Engineer for Frontier AI Code Agents
Remote Data Engineer for Frontier AI Code Agents

Mercor • Greater London

Remote
GBP 102,000 - 205,000
Frontier AI Data Engineer & Model Evaluator
Frontier AI Data Engineer & Model Evaluator

Obsidian • Greater London

Remote
GBP 327,000 - 491,000
Data Engineer for Frontier AI Model Evaluation
Data Engineer for Frontier AI Model Evaluation

Obsidian • Greater London

On-site
GBP 122,000 - 408,000
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • Greater London

On-site
GBP 122,000 - 408,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Greater London

On-site
GBP 136,000 - 205,000
AI Model Evaluation Engineer - Data Pipelines & ETL
AI Model Evaluation Engineer - Data Pipelines & ETL

Mercor • Greater London

On-site
GBP 136,000 - 204,000
Remote ML Engineer — Production AI & MLOps (Contract)
Remote ML Engineer — Production AI & MLOps (Contract)

Mercor • United Kingdom

On-site
GBP 73,696 - 100,308
Remote work
Frontier AI Agents Engineer
Frontier AI Agents Engineer

United States Digital Space LLC • Greater London

Hybrid
GBP 90,000 - 130,000