DevOps Engineer - AI Model Evaluator

Mercor

Greater London

On-site

GBP 136,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project and evaluate frontier AI coding models through structured assessments.

The work centers on realistic infrastructure engineering workflows and model evaluation, with sprint-based timeframes and a focus on reliability and scalability of cloud platforms.

Qualifications

  • 2+ years of DevOps/SRE/Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, observability tooling.
  • Regular use of AI coding agents like Cursor, Claude Code, Codex, Windsurf, Gemini CLI.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD pipelines, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps/SRE
Cloud engineering
AWS
Azure
GCP
Kubernetes
Terraform
CI/CD pipelines
Observability tooling
AI coding agents

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic infrastructure engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • Greater London

On-site
GBP 122,000 - 408,000
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Mercor • Greater London

On-site
GBP 184,000 - 255,000
Data Engineer for Frontier AI Model Evaluation
Data Engineer for Frontier AI Model Evaluation

Obsidian • Greater London

On-site
GBP 122,000 - 408,000
Frontier AI Data Engineer & Model Evaluator
Frontier AI Data Engineer & Model Evaluator

Obsidian • Greater London

Remote
GBP 327,000 - 491,000
Remote Data Engineer for Frontier AI Code Agents
Remote Data Engineer for Frontier AI Code Agents

Mercor • Greater London

Remote
GBP 102,000 - 205,000
Frontier ML Engineer — AI Coding Model Evaluator (Contract)
Frontier ML Engineer — AI Coding Model Evaluator (Contract)

Obsidian • Greater London

Remote
GBP 136,000 - 204,000
ML Engineer: Frontier AI Coding & Model Evaluation
ML Engineer: Frontier AI Coding & Model Evaluation

Obsidian • Greater London

On-site
GBP 408,000 - 612,000
AI Model Evaluation Engineer - Data Pipelines & ETL
AI Model Evaluation Engineer - Data Pipelines & ETL

Mercor • Greater London

On-site
GBP 136,000 - 204,000
Member of Technical Staff - AI
Member of Technical Staff - AI

Model ML • Greater London

On-site
GBP 90,000 - 120,000
Software Engineer, Codex Core Agents
Software Engineer, Codex Core Agents

OpenAI • Greater London

On-site
GBP 120,000 - 180,000