DevOps Engineer - AI Model Evaluator

Mercor

Berlin

Vor Ort

EUR 15.000 - 31.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments.

You will review model-generated implementations involving cloud platforms, Kubernetes, CI/CD, observability, and infrastructure automation, identifying bugs, edge cases, reliability issues, and failure modes.

Qualifikationen

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Aufgaben

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Kenntnisse

DevOps
SRE
Cloud engineering
Kubernetes
Terraform
CI/CD pipelines
Observability tooling
AWS
Azure
GCP
AI coding agents

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Jobbeschreibung

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic infrastructure engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2–3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

DevOps / SRE / Cloud Engineer (Coding Agent Experience)
DevOps / SRE / Cloud Engineer (Coding Agent Experience)

aitrainer • Deutschland

Vor Ort
EUR 43.066 - 68.906
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Mercor • Berlin

Vor Ort
EUR 393.000 - 560.000
Security Engineer - Fully Remote | Upto $85/hr
Security Engineer - Fully Remote | Upto $85/hr

Obsidian • Berlin

Remote
EUR 242.493 - 363.740
Risk Engineer - Fully Remote | Upto $80/hr
Risk Engineer - Fully Remote | Upto $80/hr

Obsidian • Berlin

Remote
EUR 242.493 - 484.986
Frontier Engineer (M/F/D)
Frontier Engineer (M/F/D)

Cognizant • Karlsruhe

Hybrid
EUR 90.000 - 130.000
Senior Software Engineer - Agent Evaluation
Senior Software Engineer - Agent Evaluation

aitrainer • Deutschland

Vor Ort
EUR 47.779 - 71.669
Senior AI Agent Evaluation Engineer
Senior AI Agent Evaluation Engineer

aitrainer • Deutschland

Vor Ort
EUR 34.453 - 60.293
Senior AI Engineer - Agentic AI Evaluation
Senior AI Engineer - Agentic AI Evaluation

Resaro AI • München

Vor Ort
EUR 90.000 - 130.000
Python Engineer, AI Coding Agent Evaluator
Python Engineer, AI Coding Agent Evaluator

g2i • Deutschland

Vor Ort
EUR 119.000 - 238.000
Forward Deployed Engineer (German-speaking)
Forward Deployed Engineer (German-speaking)

5U AI • München

Vor Ort
EUR 90.000 - 130.000