Software Engineering Evaluation Specialist

Mindrift

Doha

On-site

QAR 150,000 - 201,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mindrift is seeking a seasoned backend-focused developer task designer. You will craft realistic coding challenges in Docker, assemble self-contained task packages, and write deterministic pytest-based tests to verify fixes. Deliverables include broken code, tests, instructions, and a reference solution.

This is project-based, not permanent employment. You will design tasks that resemble real-world developer work, ensuring reproducible environments, clear Jira-like tickets, and scalable

Qualifications

  • 3+ years of production software development in one backend stack.
  • Python + pytest fluency — required regardless of primary stack.
  • Experience with Docker, Linux and Bash in containerized environments.

Responsibilities

  • Invent realistic developer scenarios — real bugs, missing features, or broken ETL.
  • Build reproducible Docker environments with pinned dependencies.
  • Write a pytest-based verification that is deterministic and non-flaky.
  • Create an instruction.md that reads like a Jira ticket for a developer.
  • Provide a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so state‑of‑the‑art agents solve the task 20–60% of the time.
  • Review other authors' tasks as a QA reviewer later.

Skills

Backend development
Python
pytest
Docker
Linux & Bash
AI tooling familiarity

Tools

Docker
pytest
CI/CD

Job description

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

About The Role

You'll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities

Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem Build a reproducible Docker environment with pinned dependencies Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix Write an instruction.md that reads like a Jira ticket a developer would receive Write a reference solve.sh proving the task is solvable Calibrate difficulty so current state-of-the-art agents solve the task 20-60% of the time Iterate based on feedback from expert QA reviewers Later: review other authors' tasks as a QA reviewer

Not in scope

Data labeling, prompt engineering Production code to ship — you design problems and verification for AI agents Leetcode puzzles — scenarios must look like real developer work Not every candidate task ships — quality over quantity

Requirements

3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it English — B2+ written

Not a fit

Data Science, ML, or Computer Vision engineers without backend-engineering output Manual QA testers without automation or test authoring Frontend-only, low-code / no-code, IT Support, or Business Analysts Engineers who have never written pytest from scratch Junior, intern, or assistant as the most recent role

Preferred Qualifications

Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (Num Py / PyTorch / Sci Py), Dev Ops, or Git internals Modern Python tooling (uv, poetry, pyproject.toml) Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov) Fuzzing or property-based testing (Hypothesis) Prior contribution to agent-evaluation benchmarks or related frameworks

Process

Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.

Time commitment

Onboarding: :10 hours per first task Steady state: :5 hours per task, 2-4 parallel tasks per author Realistic weekly load: 8-20 hours. Higher volume available for top performers You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria

Compensation

Paid contributions, rates up to $35/hour*Task-based compensation equivalent to hourly rate, depending on performance and volume Some projects include incentive payments Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

Mindrift • Doha

On-site
QAR 150,000 - 251,000
Mindrift project-based engagement
Senior Python Engineer - AI Task Designer (Project-Based)
Senior Python Engineer - AI Task Designer (Project-Based)

Mindrift • Doha

On-site
QAR 301,000 - 1,003,000
Paid per task
AI Evaluation Engineer: Build Real-World Coding Agent Tests
AI Evaluation Engineer: Build Real-World Coding Agent Tests

Mindrift • Doha

On-site
QAR 150,000 - 251,000
Mindrift project-based engagement
Contract AI Evaluation Engineer - Design & Test Agent Tasks
Contract AI Evaluation Engineer - Design & Test Agent Tasks

Employment • Doha

On-site
QAR 150,000 - 251,000
AI Agent Evaluation Engineer
AI Agent Evaluation Engineer

Mindrift • Doha

On-site
QAR 207,000 - 295,000
AI Task Architect for Backend Evaluation
AI Task Architect for Backend Evaluation

Mindrift • Doha

On-site
QAR 150,000 - 201,000
Compliance Analyst
Compliance Analyst

Employment • Doha

On-site
QAR 150,000 - 251,000
Machine Learning Engineer
Machine Learning Engineer

Employment • Doha

On-site
QAR 437,000 - 655,000
AI Software Engineer - Full Stack Developer (Java) at EPAM Systems
AI Software Engineer - Full Stack Developer (Java) at EPAM Systems

EPAM Systems • Doha

On-site
QAR 180,000 - 300,000
Healthcare & life insurance
End of service gratuity
Air travel tickets for expatriates
+2
Remote Full-Stack Web App Developer (AI Pilot)
Remote Full-Stack Web App Developer (AI Pilot)

Employment • Doha

On-site
QAR 248,000 - 354,000