AI-Driven DevOps Engineer & Model Evaluator

Mercor

San Francisco (CA)

On-site

USD 15,000 - 22,000

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and CI/CD tools.

This sprint-based role involves evaluating bugs, edge cases, and failure modes with professional engineering judgment, applying real-world reliability concepts to production-scale systems.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
SRE
Cloud engineering
Observability
AI coding agents experience
Infrastructure as code

Tools

AWS
Azure
GCP
Kubernetes
Terraform
CI/CD pipelines
Observability tooling

Job description

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and CI/CD tools.

This sprint-based role involves evaluating bugs, edge cases, and failure modes with professional engineering judgment, applying real-world reliability concepts to production-scale systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI-Powered Data Engineer & Model Evaluator
AI-Powered Data Engineer & Model Evaluator

Mercor • New York (NY)

On-site
USD 38,000 - 46,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Obsidian • Miami (FL)

On-site
USD 455,000 - 647,000
Pay per task
Frontier AI Code Engineer
Frontier AI Code Engineer

Mercor • New York (NY)

On-site
USD 220,000 - 551,000
Frontier AI Data Engineer: Model Evaluation & Pipelines
Frontier AI Data Engineer: Model Evaluation & Pipelines

Obsidian • Philadelphia

On-site
USD 455,000 - 647,000
Frontier AI Data Engineer: ETL & Model Evaluation
Frontier AI Data Engineer: ETL & Model Evaluation

Mercor • Philadelphia

On-site
USD 179,000 - 276,000
Frontier AI Infrastructure Engineer (Contract)
Frontier AI Infrastructure Engineer (Contract)

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000