AI-Driven DevOps Evaluator for Frontier Coding Agents

Mercor

New York (NY)

On-site

USD 441,000 - 661,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments.

The work focuses on realistic infrastructure engineering workflows and model evaluation, reviewing model-generated implementations across cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.

Qualifications

  • : 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • : Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • : Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • : Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • : Experience supporting production-scale systems is preferred.

Responsibilities

  • : Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • : Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • : Identify bugs, edge cases, reliability issues, and failure modes.
  • : Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • : Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud engineering
AWS
Azure
GCP
Kubernetes
Terraform
CI/CD pipelines
Observability tooling

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments.

The work focuses on realistic infrastructure engineering workflows and model evaluation, reviewing model-generated implementations across cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven DevOps Evaluator for Frontier Code Models
AI-Driven DevOps Evaluator for Frontier Code Models

Mercor • San Francisco (CA)

Hybrid
USD 55,000 - 91,000
AI-Driven DevOps Engineer & Model Evaluator
AI-Driven DevOps Engineer & Model Evaluator

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
Frontier AI Code Engineer — ML Systems & Evaluation
Frontier AI Code Engineer — ML Systems & Evaluation

Mercor • New York (NY)

On-site
USD 179,000 - 276,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Frontier AI Data Engineer: ETL & Model Evaluation
Frontier AI Data Engineer: ETL & Model Evaluation

Mercor • Philadelphia

On-site
USD 179,000 - 276,000
Frontier AI Infrastructure Engineer (Contract)
Frontier AI Infrastructure Engineer (Contract)

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000
Frontier AI Data Engineer — Model Evaluation & ETL
Frontier AI Data Engineer — Model Evaluation & ETL

Mercor • San Francisco (CA)

On-site
USD 207,000 - 234,000
AI Code-Agent Evaluator: Frontier DevOps Engineer
AI Code-Agent Evaluator: Frontier DevOps Engineer

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • New York (NY)

On-site
USD 441,000 - 661,000