DevOps Engineer - AI Model Evaluator

mercor

Italia

Remote

EUR 88,000 - 120,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is seeking a DevOps / SRE / Cloud Engineer (Coding Agent Experience) for a remote contract role. You will use frontier AI coding agents to tackle infrastructure tasks across cloud platforms, Kubernetes, CI/CD, observability, and infrastructure automation, while evaluating model outputs for reliability.

Responsibilities include reviewing model-generated implementations, identifying edge cases and failure modes, and comparing outputs from multiple frontier models to guide engineering

Qualifications

  • 2+ years of professional DevOps, SRE, or cloud engineering experience.
  • Experience with AWS, Azure, or GCP and Kubernetes.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability solutions.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD, observability, and infra automation.
  • Identify bugs, edge cases, reliability issues, and failure modes in model outputs.
  • Compare outputs from multiple frontier models to assess strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud Engineering

Tools

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD
Observability tooling

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position

DevOps / SRE / Cloud Engineer (Coding Agent Experience)

Type

Contract

Compensation

$85/hour

Location

Remote

Role Responsibilities
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , observability , and infrastructure automation .
  • Identify bugs, edge cases, reliability issues, and failure modes in model outputs.
  • Compare outputs from multiple frontier models to assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Qualifications Must-Have
  • 2+ years of professional DevOps , SRE , or Cloud Engineering experience.
  • Experience with AWS , Azure , GCP , Kubernetes , Terraform , CI/CD pipelines , or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
Preferred
  • Experience supporting production-scale systems.
Compensation & Legal

$400 per accepted task

Compensation tied to accepted work.

Application Process (Takes 20–30 mins to complete)
  • Upload resume
  • AI interview based on your resume
  • Submit form
Resources & Support
  • For details about the interview process and platform information, please check:
  • For any help or support, reach out to:

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote DevOps Engineer for AI Model Evaluation
Remote DevOps Engineer for AI Model Evaluation

mercor • Italy

Remote
EUR 88,000 - 120,000
Freelance Agent Evaluation Engineer
Freelance Agent Evaluation Engineer

Mindrift • Milano

On-site
EUR 30,000 - 48,000
Admin Support Specialist - Fully Remote | Upto $80/hr
Admin Support Specialist - Fully Remote | Upto $80/hr

mercor • Italy

Remote
EUR 37,000 - 98,000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Milano

On-site
EUR 23,623 - 35,435
Flexible, remote work
Competitive pay up to $30/hour
Gain experience in advanced AI projects
Generalist Expert - Content Evaluator
Generalist Expert - Content Evaluator

Mercor • Lazio

Remote
EUR 60,000 - 84,000
AI Engineer/Agentic Workflow Platform
AI Engineer/Agentic Workflow Platform

NTT DATA Europe & Latam • Emilia-Romagna

On-site
EUR 80,000 - 120,000
Program Manager - Fully Remote | Upto $160/hr
Program Manager - Fully Remote | Upto $160/hr

mercor • Italy

Remote
EUR 98,000 - 197,000
Academic Evaluator - Fully Remote | Upto $160/hr
Academic Evaluator - Fully Remote | Upto $160/hr

mercor • Italy

Remote
EUR 98,000 - 197,000
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Roma

On-site
EUR 32,861 - 46,802
Competitive pay rates up to $34/hour
Flexible freelance work hours
Experience on advanced AI projects
Lead Engineer, AI Platform
Lead Engineer, AI Platform

Jobgether SRL • Italy

On-site
EUR 136,000 - 167,000
Annual $170,000 USD cash compensation
Equity with refresh grants
Fully remote work environment
+2