Senior AI Evaluation Engineer – Agent Testing & Tasks

Dorado

United States

Remote

USD 55,104 - 82,656

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift seeks contributors to build a dataset for evaluating AI coding agents. You will create tasks and evaluation criteria within realistic simulated environments, including a virtual codebase, infrastructure, tickets, and conversations to form a believable development history.

You will design prompts, define success criteria, and write tests that accept diverse correct solutions while rejecting incorrect ones. This is a project-based engagement rather than permanent employment.

Qualifications

  • 5+ years in software development.
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis.
  • Experience writing tests (functional, integration).
  • English proficiency - B2+.

Responsibilities

  • Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations).
  • Design tasks from intermediate states of these environments — craft the prompt, define what “solved” means, and ensure solvable by an AI agent.
  • Write tests that verify agent solutions — accept all valid approaches and reject incorrect ones, not too strict or lenient.
  • Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until evaluation is fair and robust.

Skills

5+ years in software development
Python
JavaScript/TypeScript
Docker
Postgres
Kafka
Redis
Testing (functional, integration)
English proficiency (B2+)

Tools

FastAPI
React
Docker
PostgreSQL
Kafka
Redis

Job description

Mindrift seeks contributors to build a dataset for evaluating AI coding agents. You will create tasks and evaluation criteria within realistic simulated environments, including a virtual codebase, infrastructure, tickets, and conversations to form a believable development history.

You will design prompts, define success criteria, and write tests that accept diverse correct solutions while rejecting incorrect ones. This is a project-based engagement rather than permanent employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Agent Evaluation Engineer
Senior AI Agent Evaluation Engineer

Worky • United States

Remote
USD 55,000 - 69,000
AI Evaluation Engineer — Flexible Hours, High-Impact Testing
AI Evaluation Engineer — Flexible Hours, High-Impact Testing

Mindrift • United States

On-site
USD 41,000 - 69,000
Senior Python Engineer — AI Task Designer & Evaluator
Senior Python Engineer — AI Task Designer & Evaluator

Dorado • United States

Remote
Software Engineer - AI Agent Evaluation
Software Engineer - AI Agent Evaluation

aitrainer • New York (NY), Northern (KY)

Hybrid
USD 57,000 - 81,000
Python Engineer - AI Coding Agent Tests (Project)
Python Engineer - AI Coding Agent Tests (Project)

Dorado • United States

Remote
AI Coding Evaluator & Testing Architect (Python)
AI Coding Evaluator & Testing Architect (Python)

Mindrift • United States

On-site
USD 55,000 - 69,000
Backend Engineer: AI Task Author & QA Specialist
Backend Engineer: AI Task Author & QA Specialist

Dorado • United States

Remote
USD 39,000 - 48,000
Software Engineer - AI Agent Evaluation (Remote) at Toloka AI
Software Engineer - AI Agent Evaluation (Remote) at Toloka AI

aitrainer • New York (NY), Northern (KY)

Hybrid
USD 57,000 - 81,000
AI Evaluation Engineer: Coding Task Architect
AI Evaluation Engineer: Coding Task Architect

United States Digital Space LLC • United States

Remote
USD 55,000 - 69,000
Remote AI Evaluation Engineer for Code Tasks
Remote AI Evaluation Engineer for Code Tasks

YO AI Labs • California (MO)

Remote
USD 50,000 - 80,000