DevOps Engineer - AI Model Evaluator

Mercor

Toronto

On-site

CAD 18,000 - 27,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project, evaluating and improving frontier AI coding models through structured technical assessments. You will work on realistic infrastructure engineering workflows and model evaluation in a fast-paced sprint environment.

The role focuses on reviewing model-generated infrastructure, identifying issues, and applying engineering judgment to complex tasks across cloud, Kubernetes, CI/CD, observability, and

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud Engineering
Observability
AI coding agents

Tools

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD pipelines

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic infrastructure engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer - AI Model Evaluator
Data Engineer - AI Model Evaluator

Mercor • Toronto

On-site
CAD 193,000 - 248,000
Data Engineer - AI Model Evaluator
Data Engineer - AI Model Evaluator

Obsidian • Toronto

On-site
CAD 413,000 - 689,000
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • Toronto

On-site
CAD 634,000 - 903,000
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Obsidian • Toronto

On-site
CAD 767,000 - 863,000
Frontier AI Data Engineer for Model Evaluation
Frontier AI Data Engineer for Model Evaluation

Obsidian • Toronto

Remote
CAD 227,000 - 324,000
Remote Rust Developer
Remote Rust Developer

turing • Canada

Remote
CAD 78,000 - 118,000
Remote Software Developer
Remote Software Developer

turing • Canada

Remote
CAD 98,000 - 137,000
Forward Deployed Engineer
Forward Deployed Engineer

Robots & Pencils LP • Canada

On-site
CAD 176,612 - 243,680
Forward Deployed AI Engineer
Forward Deployed AI Engineer

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Kinvie • Toronto

On-site
CAD 90,000 - 120,000