AI Code Agent Evaluator for DevOps & Infra

Mercor

Warszawa

On-site

PLN 1,794,000 - 2,307,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on realistic infrastructure engineering workflows and model evaluation.

The role requires 2+ years of DevOps/SRE/Cloud Engineering experience and hands-on work with AWS/Azure/GCP, Kubernetes, Terraform, CI/CD, and observability tools; familiarity with AI coding agents is a plus.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud engineering
AI coding agents

Tools

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD tooling
Observability tooling

Job description

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on realistic infrastructure engineering workflows and model evaluation.

The role requires 2+ years of DevOps/SRE/Cloud Engineering experience and hands-on work with AWS/Azure/GCP, Kubernetes, Terraform, CI/CD, and observability tools; familiarity with AI coding agents is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier Cloud Engineer for AI-Driven Infra Evaluation
Frontier Cloud Engineer for AI-Driven Infra Evaluation

Mercor • Warszawa

Remote
PLN 618,000 - 1,029,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Warszawa

On-site
PLN 1,794,000 - 2,307,000
Cloud Engineer - Fully Remote
Cloud Engineer - Fully Remote

Mercor • Warszawa

Remote
PLN 618,000 - 1,029,000
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Mercor • Warszawa

On-site
PLN 120,000 - 190,000
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Warszawa

Remote
PLN 180,000 - 240,000
Frontier AI Safety Red Team Lead
Frontier AI Safety Red Team Lead

Mercor • Warszawa

On-site
PLN 180,000 - 300,000
Senior AI Engineer
Senior AI Engineer

Bayer CropScience Limited • Warszawa

On-site
PLN 100,000 - 140,000
Senior AI Platform Engineer — Scale AI Agents in Cloud (Remote)
Senior AI Platform Engineer — Scale AI Agents in Cloud (Remote)

EPAM Systems • Poland

Hybrid
PLN 180,000 - 300,000
Hybrid work model
Remote in Poland
Work abroad availability
+2
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Warszawa

Remote
PLN 200,000 - 320,000
AI Coding Agent Evaluator (Project-Based)
AI Coding Agent Evaluator (Project-Based)

Mindrift • Poland

On-site