DevOps Engineer - AI Model Evaluator

Mercor

Paris

Sur place

EUR 406 000 - 549 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments focused on realistic infrastructure engineering workflows.

The role involves reviewing model-generated implementations across cloud platforms, Kubernetes, CI/CD, observability, and automation. You will identify bugs, edge cases, and failure modes while applying sound engineering judgment.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsabilités

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Connaissances

DevOps/SRE
Cloud engineering
AI coding agents
Model evaluation
Production systems

Outils

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD pipelines
Observability tooling

Description du poste

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic infrastructure engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Lacaussade

Sur place
EUR 84 000 - 119 000
AI DevOps Evaluator for Frontier Code Agents
AI DevOps Evaluator for Frontier Code Agents

Mercor • Paris

Sur place
EUR 406 000 - 549 000
Remote DevOps Engineer for AI Model Evaluation
Remote DevOps Engineer for AI Model Evaluation

Mercor • Lacaussade

Sur place
EUR 84 000 - 119 000
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift • France

Sur place
DevOps / MLOps Engineer
DevOps / MLOps Engineer

Jobtailor • Strasbourg

Sur place
EUR 70 000 - 110 000
Forward Deployed Engineer
Forward Deployed Engineer

Axiōma Search • Paris

Hybride
EUR 70 000 - 110 000
DevOps / IaC Engineer
DevOps / IaC Engineer

ixolabs.ai • France

À distance
EUR 60 000 - 90 000
Freelance Agent Evaluation Analyst
Freelance Agent Evaluation Analyst

Mindrift • Paris

À distance
Flexible working hours
Competitive pay
Remote work options
+1
DevOps - Continuous Integration
DevOps - Continuous Integration

Linuxcareers • Paris

Sur place
EUR 45 000 - 70 000
Applied AI Engineer
Applied AI Engineer

Norbert Health • Paris

Sur place
EUR 70 000 - 90 000
Competitive salary and equity
High autonomy and technical ownership
Transparent, mission-driven culture