AI Frontier Code Agent Evaluator for DevOps and Infra

Mercor

Madrid

Presencial

EUR 179.000 - 239.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on infrastructure engineering workflows and model evaluation.

If you have 2+ years in DevOps/SRE, hands-on cloud and Kubernetes, and enjoy working with AI coding agents like Cursor or Claude Code, join us to assess and enhance production-scale systems.

Formación

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsabilidades

  • Use frontier AI coding agents to complete and evaluate infrastructure tasks.
  • Review model-generated implementations across cloud, Kubernetes, CI/CD, observability, and automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure scenarios.

Conocimientos

DevOps
SRE
Cloud engineering

Herramientas

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD
Observability tooling

Descripción del empleo

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments, focusing on infrastructure engineering workflows and model evaluation.

If you have 2+ years in DevOps/SRE, hands-on cloud and Kubernetes, and enjoy working with AI coding agents like Cursor or Claude Code, join us to assess and enhance production-scale systems.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Madrid

Presencial
EUR 179.000 - 239.000
Remote DevOps Engineer for AI Model Evaluation
Remote DevOps Engineer for AI Model Evaluation

Mercor • Madrid

Presencial
EUR 147.000
Frontier AI Safety Evaluator & Alignment Expert
Frontier AI Safety Evaluator & Alignment Expert

Mercor • Madrid

Presencial
EUR 70.000 - 110.000
Frontier AI Safety & Evaluation Specialist
Frontier AI Safety & Evaluation Specialist

Mercor • Madrid

A distancia
EUR 60.000 - 90.000
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Madrid

A distancia
EUR 100.000 - 140.000
Frontier AI Safety Red Team Specialist
Frontier AI Safety Red Team Specialist

Mercor • Madrid

A distancia
EUR 60.000 - 100.000
Frontier AI Safety Evaluator & Policy Feedback Lead
Frontier AI Safety Evaluator & Policy Feedback Lead

Mercor • Madrid

Presencial
EUR 55.000 - 90.000
AI Coding Agent Evaluator - Design & QA Real-World Tasks
AI Coding Agent Evaluator - Design & QA Real-World Tasks

Mindrift • Madrid

Presencial
Agentic AI Engineer
Agentic AI Engineer

Mondia • Madrid

Híbrido
EUR 65.000 - 90.000
Hybrid work
Company bonus
Private health insurance
+2
AI Developer
AI Developer

Avenue Code • España

Presencial
EUR 90.000 - 140.000