DevOps Engineer - AI Model Evaluator

Mercor

Warszawa

On-site

PLN 1,794,000 - 2,307,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on realistic infrastructure engineering workflows and model evaluation.

The role requires 2+ years of DevOps/SRE/Cloud Engineering experience and hands-on work with AWS/Azure/GCP, Kubernetes, Terraform, CI/CD, and observability tools; familiarity with AI coding agents is a plus.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud engineering
AI coding agents

Tools

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD tooling
Observability tooling

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic infrastructure engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer - Fully Remote
Cloud Engineer - Fully Remote

Mercor • Warszawa

Remote
PLN 618,000 - 1,029,000
AI Code Agent Evaluator for DevOps & Infra
AI Code Agent Evaluator for DevOps & Infra

Mercor • Warszawa

On-site
PLN 1,794,000 - 2,307,000
Frontier Cloud Engineer for AI-Driven Infra Evaluation
Frontier Cloud Engineer for AI-Driven Infra Evaluation

Mercor • Warszawa

Remote
PLN 618,000 - 1,029,000
Senior Software Engineer - AI Model Evaluation
Senior Software Engineer - AI Model Evaluation

Mindrift • Poland

On-site
Senior Software Engineer - AI Coding Benchmarks
Senior Software Engineer - AI Coding Benchmarks

Mindrift • Poland

On-site
DevOps Engineer
DevOps Engineer

Antler • Województwo wielkopolskie

On-site
PLN 50,000 - 70,000
DevOps Engineer
DevOps Engineer

Antler • Warszawa

On-site
PLN 70,000 - 100,000
Senior Software Engineer - Testing AI Coding Agents
Senior Software Engineer - Testing AI Coding Agents

Mindrift • Poland

On-site
DevOps Engineer
DevOps Engineer

Duku AI • Łódź

On-site
PLN 80,000 - 100,000
DevOps Engineer
DevOps Engineer

Antler • Kraków

On-site
PLN 80,000 - 110,000