AI-Driven DevOps Engineer & Model Evaluator

Obsidian

San Francisco (CA)

On-site

USD 551,040 - 826,560

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Obsidian is seeking contributors for an innovative Frontier Code Agents project that evaluates and improves AI coding models. The role requires 2+ years of experience in DevOps, SRE, or Cloud Engineering and proficiency with cloud platforms like AWS, Azure, and GCP.

You will utilize AI coding agents to handle complex engineering tasks and assess their outputs across multiple models. The compensation is $400 per accepted task, making this a great opportunity for professionals to apply their skills in real-world scenarios.

Qualifications

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience.
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Identify bugs, edge cases, reliability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Skills

DevOps
Cloud Engineering
Kubernetes
CI/CD pipelines
Observability tooling

Tools

AWS
Azure
GCP
Terraform
Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

Obsidian is seeking contributors for an innovative Frontier Code Agents project that evaluates and improves AI coding models. The role requires 2+ years of experience in DevOps, SRE, or Cloud Engineering and proficiency with cloud platforms like AWS, Azure, and GCP.

You will utilize AI coding agents to handle complex engineering tasks and assess their outputs across multiple models. The compensation is $400 per accepted task, making this a great opportunity for professionals to apply their skills in real-world scenarios.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Code-Agent Evaluator: Frontier DevOps Engineer
AI Code-Agent Evaluator: Frontier DevOps Engineer

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
Frontier Cloud Engineer: AI Coding Agents
Frontier Cloud Engineer: AI Coding Agents

Obsidian • New York (NY)

Remote
USD 551,000 - 1,102,000
AI-Driven Data Engineer - Frontier Pipeline Evaluator
AI-Driven Data Engineer - Frontier Pipeline Evaluator

Obsidian • New York (NY)

Remote
USD 193,000 - 248,000
AI Systems Engineer: Frontier Model Evaluator
AI Systems Engineer: Frontier Model Evaluator

Obsidian • New York (NY)

On-site
USD 200,000
Frontier ML Engineer: AI Coding Agent Evaluator
Frontier ML Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

Hybrid
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • San Francisco (CA)

Hybrid
USD 55,000 - 91,000
AI-Driven DevOps Evaluator for Frontier Code Models
AI-Driven DevOps Evaluator for Frontier Code Models

Mercor • San Francisco (CA)

Hybrid
USD 55,000 - 91,000
Backend Engineer: AI Coding Agent Evaluator
Backend Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

On-site
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000