AI DevOps Evaluator for Frontier Code Agents

Mercor

Stockholms kommun

On-site

SEK 4,723,000 - 5,773,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mercor partners with a leading AI research lab on a Frontier Code Agents project, focusing on infrastructure engineering workflows and model evaluation. Contributors evaluate model-generated code and assess reliability, edge cases, and failure modes across complex cloud-based systems.

You will compare outputs from frontier models, exercise professional engineering judgment, and help improve frontier AI coding agents in realistic scenarios.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud engineering
AWS
Azure
GCP
Kubernetes
Terraform
CI/CD
Observability
AI coding agents

Tools

Kubernetes
Terraform
CI/CD pipelines

Job description

Mercor partners with a leading AI research lab on a Frontier Code Agents project, focusing on infrastructure engineering workflows and model evaluation. Contributors evaluate model-generated code and assess reliability, edge cases, and failure modes across complex cloud-based systems.

You will compare outputs from frontier models, exercise professional engineering judgment, and help improve frontier AI coding agents in realistic scenarios.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Stockholms kommun

On-site
SEK 4,723,000 - 5,773,000
Frontier AI Safety Evaluator - Expert Quality Review
Frontier AI Safety Evaluator - Expert Quality Review

Mercor • Stockholms kommun

On-site
SEK 650,000 - 900,000
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Mercor • Stockholms kommun

Remote
SEK 700,000 - 1,100,000
Frontier AI Safety Evaluator & Policy Feedback Specialist
Frontier AI Safety Evaluator & Policy Feedback Specialist

Mercor • Stockholms kommun

On-site
SEK 780,000 - 1,050,000
Senior Frontier AI Safety Red Team Expert
Senior Frontier AI Safety Red Team Expert

Mercor • Stockholms kommun

On-site
SEK 1,100,000 - 1,500,000
Python Engineer - Freelance AI Trainer
Python Engineer - Freelance AI Trainer

Mindrift • Sweden

On-site
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Stockholms kommun

Remote
SEK 700,000 - 1,000,000
AI Content Evaluator (UK/Europe): Elevate Real-World Docs
AI Content Evaluator (UK/Europe): Elevate Real-World Docs

Mercor • Stockholms kommun

On-site
SEK 555,000 - 777,000
Senior Python Engineer - AI Task Designer & Evaluator
Senior Python Engineer - AI Task Designer & Evaluator

Mindrift • Sweden

On-site
AI Document Quality Assessor (UK/Europe)
AI Document Quality Assessor (UK/Europe)

Mercor • Stockholms kommun

Remote
SEK 381,000 - 685,000