DevOps Engineer & AI Model Evaluator for Frontier Coding

Mercor

Oslo

On-site

NOK 4,448,000 - 6,018,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors evaluate frontier AI coding models through structured technical assessments focused on infrastructure engineering workflows and model evaluation.

Role emphasizes reviewing cloud platforms, Kubernetes, CI/CD, observability, and automation while applying engineering judgment to realistic scenarios. Sprint-based work runs 12–24 hour stretches with compensation per accepted task.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps experience
Cloud engineering
SRE
Kubernetes
Terraform
CI/CD pipelines
Observability tooling
AI coding agents
Production-scale systems

Tools

AWS
Azure
GCP

Job description

Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors evaluate frontier AI coding models through structured technical assessments focused on infrastructure engineering workflows and model evaluation.

Role emphasizes reviewing cloud platforms, Kubernetes, CI/CD, observability, and automation while applying engineering judgment to realistic scenarios. Sprint-based work runs 12–24 hour stretches with compensation per accepted task.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Oslo

On-site
NOK 4,448,000 - 6,018,000
Frontier AI Safety Red Team Specialist
Frontier AI Safety Red Team Specialist

Mercor • Oslo

On-site
NOK 900,000 - 1,400,000
Frontier AI Safety Evaluator & Model Alignment
Frontier AI Safety Evaluator & Model Alignment

Mercor • Oslo

On-site
NOK 900,000 - 1,300,000
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Oslo

Remote
NOK 1,000,000 - 1,400,000
AI Safety Evaluator: Frontier Models & Policy Alignment
AI Safety Evaluator: Frontier Models & Policy Alignment

Obsidian • Oslo

On-site
NOK 700,000 - 1,100,000
Forward Deployed Engineer
Forward Deployed Engineer

Wonderful • Norway

On-site
NOK 900,000 - 1,200,000
Senior Ops Lead - High-Impact AI Data Projects (Equity)
Senior Ops Lead - High-Impact AI Data Projects (Equity)

Mercor • London

On-site
NOK 1,129,000 - 1,882,000
Proximity bonus
Meals stipend
Gym membership
+4
Senior DevOps Engineer
Senior DevOps Engineer

Newcode.ai • Oslo

Hybrid
NOK 746,000 - 1,120,000
Collaborative team environment
Flexible working conditions
Impact on AI product development
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Obsidian • Oslo

Remote
NOK 855,000 - 1,330,000
AI Engineer
AI Engineer

Appfarm • Oslo

On-site
NOK 900,000 - 1,300,000