Data Engineer - AI Model Evaluation

Mercor

San Francisco (CA)

On-site

USD 207,000 - 234,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models through structured technical assessments and focus on realistic data engineering workflows.

The role involves reviewing model-generated ETL pipelines, data warehouses, analytics platforms, and distributed data systems, and applying engineering judgment to identify bugs and scalability issues. Sprint-based work with rapid task turnover is expected.

Qualifications

  • 2+ years of professional data engineering experience.
  • Experience building ETL pipelines, data warehouses, analytics platforms, or distributed data systems.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated data infrastructure and pipeline implementations.
  • Experience operating large-scale data platforms is preferred.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex data engineering tasks.
  • Review model-generated implementations involving ETL pipelines, data warehouses, analytics platforms, and distributed data systems.
  • Identify bugs, edge cases, scalability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic data engineering scenarios.

Skills

Data engineering
ETL pipelines
Data warehouses
Analytics platforms
Distributed data systems
AI coding agents

Tools

Cursor
Claude Code
Codex
Windsurf
Gemini CLI

Job description

About the Role
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project.
  • Contributors help evaluate and improve frontier AI coding models through structured technical assessments.
  • The work focuses on realistic data engineering workflows and model evaluation.
  • Spots are limited and filling quickly on a first come, first serve basis.
What You'll Do
  • Use frontier AI coding agents to complete and evaluate complex data engineering tasks.
  • Review model-generated implementations involving ETL pipelines, data warehouses, analytics platforms, and distributed data systems.
  • Identify bugs, edge cases, scalability issues, and failure modes.
  • Compare outputs from multiple frontier models and assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic data engineering scenarios.
Time Commitment
  • Sprint based project that runs in 12-24 hour stretches based on client requirement.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2-3 hours after ramp-up.
  • Compensation is tied to accepted work.
Who Should Apply
  • 2+ years of professional data engineering experience.
  • Experience building ETL pipelines, data warehouses, analytics platforms, or distributed data systems.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated data infrastructure and pipeline implementations.
  • Experience operating large-scale data platforms is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • San Francisco (CA)

On-site
USD 207,000 - 551,000
Data Engineer - AI Model Evaluator - AI Trainer
Data Engineer - AI Model Evaluator - AI Trainer

Mercor • Philadelphia

On-site
USD 179,000 - 276,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • New York (NY)

On-site
USD 441,000 - 661,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • San Francisco (CA)

Hybrid
USD 55,000 - 91,000
DevOps Engineer - AI Model Evaluator - AI Trainer
DevOps Engineer - AI Model Evaluator - AI Trainer

Obsidian • Miami (FL)

On-site
USD 455,000 - 647,000
Pay per task
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Obsidian • New York (NY)

On-site
USD 455,000 - 647,000
Data Engineer - Fully Remote | Upto $80/hr
Data Engineer - Fully Remote | Upto $80/hr

Obsidian • San Francisco (CA)

Remote
Frontier AI Data Engineer — Model Evaluation & ETL
Frontier AI Data Engineer — Model Evaluation & ETL

Mercor • San Francisco (CA)

On-site
USD 207,000 - 234,000