AI Evaluation Engineer

Mindrift

Doha

On-site

QAR 207,000 - 295,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible schedule
Competitive hourly rate

Job summary

Mindrift is assembling a dataset and evaluation framework to test AI coding agents. You will construct realistic developer environments, including codebases, infrastructure, and contextual artifacts such as tickets and docs, to simulate a believable history.

You will design tasks from intermediate states, define solvable criteria for AI agents, and craft tests that tolerate multiple valid approaches while rejecting incorrect ones.

Qualifications

  • 5+ years in software development
  • Core stack: Python (Fast API), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Responsibilities

  • Build realistic developer environments: a virtual company with codebase, infrastructure, and context
  • Design tasks from intermediate states of these environments and define what solved means
  • Write tests that verify agent solutions, accepting all valid approaches and rejecting incorrect ones
  • Iterate on tasks and tests based on QA feedback to ensure fairness and robustness

Skills

Python
JavaScript/TypeScript
React
Docker
PostgreSQL
Kafka
Redis
Testing
English proficiency

Tools

Docker
PostgreSQL
Kafka
Redis

Job description

Industry Information Technology and Services
What this opportunity involves

We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

You'll create challenging tasks and evaluation criteria within realistic simulated environments:
  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What This Is NOT
  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What We Look For
  • 5+ years in software development Core stack: Python (Fast API), Java Script/Type Script (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

Compensation

Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at :20 hours each; you set your own schedule.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Build Real-World Coding Challenges
AI Evaluation Engineer: Build Real-World Coding Challenges

Mindrift • Doha

On-site
QAR 207,000 - 295,000
Flexible schedule
Competitive hourly rate
Contract AI Evaluation Engineer - Design & Test Agent Tasks
Contract AI Evaluation Engineer - Design & Test Agent Tasks

Employment • Doha

On-site
QAR 150,000 - 251,000
Remote Evaluation Engineer for AI Task Testing (Freelance)
Remote Evaluation Engineer for AI Task Testing (Freelance)

Mindrift • Doha

On-site
QAR 150,000 - 251,000
Remote freelance project
Part-time role
Flexible hours
+1
Remote Arabic Language Expert & AI Training Evaluator
Remote Arabic Language Expert & AI Training Evaluator

YO IT Consulting • Qatar

On-site
QAR 125,372 - 250,745
AI Trainer - Remote
AI Trainer - Remote

YO AI Labs • Doha

Remote
QAR 125,000 - 251,000
Software Engineer - Open Source Contributions - Remote
Software Engineer - Open Source Contributions - Remote

YO AI Labs • Doha

Remote
QAR 218,000 - 328,000
Arabic Language Expert - Remote
Arabic Language Expert - Remote

YO IT Consulting • Qatar

On-site
QAR 125,372 - 250,745
Freelance Annotator (English) - AI Trainer
Freelance Annotator (English) - AI Trainer

Tanqeeb • Doha

Remote
QAR 50,000 - 75,000
Remote Senior Full-Stack Engineer for AI Training
Remote Senior Full-Stack Engineer for AI Training

YO AI Labs • Doha

Remote
QAR 301,000 - 451,000
AI Content Evaluator & Data Annotator (Remote)
AI Content Evaluator & Data Annotator (Remote)

YO AI Labs • Doha

Remote
QAR 125,000 - 251,000