Software Engineering Evaluation Specialist

Mindrift

Dubai

On-site

AED 126,478,000 - 177,069,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift connects specialists with project-based AI opportunities. You will design realistic coding tasks, create self-contained Docker environments, and produce tests that verify outcomes without revealing fixes. Tasks are reviewed by QA and iterated based on feedback, ensuring quality and solvability.

The role emphasizes backend-focused development, Python/pytest fluency, and Docker expertise, with project-based compensation and non-permanent engagement.

Qualifications

  • 3+ years of production software development in one backend stack (Python, Go, Node.js, Java, or Rust).
  • Python + pytest fluency required; the task harness is pytest-based even if the app is in another language.
  • Docker authoring with reproducible environments, pinned dependencies, and non-root user.
  • Linux and Bash proficiency for debugging containers (strace, lsof, journalctl).
  • AI coding agent experience with Claude Code, Cursor, Roo Code, or similar, on non-trivial work.
  • English – B2+ written.

Responsibilities

  • Invent a realistic developer scenario (real bug, broken ETL, missing feature).
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes (deterministic, non-flaky).
  • Create an instruction.md that reads like a Jira ticket.
  • Provide a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so state-of-the-art agents solve 20–60% of tasks.
  • Iterate based on QA feedback; review other authors’ tasks as QA reviewer.

Skills

Python
pytest
Docker
Linux & Bash
English (B2+)
Backend development

Tools

Docker
pytest

Job description

Job Description:

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

About the Role

You’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities:
  • Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.
  • Write an instruction.md that reads like a Jira ticket a developer would receive.
  • Write a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.
  • Iterate based on feedback from expert QA reviewers.
  • Later: review other authors’ tasks as a QA reviewer.
Not in scope
  • Data labeling, prompt engineering.
  • Production code to ship — you design problems and verification for AI agents.
  • Leetcode puzzles — scenarios must look like real developer work.
  • Not every candidate task ships — quality over quantity.
Requirements
  • 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
  • Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
  • Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
  • English — B2+ written.
Not a fit
  • Data Science, ML, or Computer Vision engineers without backend-engineering output.
  • Manual QA testers without automation or test authoring.
  • Frontend-only, low-code / no-code, IT Support, or Business Analysts.
  • Engineers who have never written pytest from scratch.
  • Junior, intern, or assistant as the most recent role.
Preferred qualifications
  • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
  • Modern Python tooling (uv, poetry, pyproject.toml).
  • Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).
  • Fuzzing or property-based testing (Hypothesis).
  • Prior contribution to agent-evaluation benchmarks or related frameworks.
Process
Time commitment
  • Onboarding: ~10 hours per first task.
  • Steady state: ~5 hours per task, 2–4 parallel tasks per author.
  • Realistic weekly load: 8–20 hours. Higher volume available for top performers.
  • You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
Compensation:
  • Paid contributions, rates up to $35/hour*.
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.

*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.

Requirements:
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Python Engineer - AI Coding Agent Evaluation Freelance
Senior Python Engineer - AI Coding Agent Evaluation Freelance

Mindrift • Dubai

On-site
AED 620,000 - 1,033,000
Senior Python Engineer - AI Task Design & Evaluation
Senior Python Engineer - AI Task Design & Evaluation

Mindrift • Dubai

On-site
AED 620,000 - 1,033,000
AI Task Designer for Backend Evaluations
AI Task Designer for Backend Evaluations

Mindrift • Dubai

On-site
AED 126,478,000 - 177,069,000
.NET Developer - Remote
.NET Developer - Remote

YO IT Consulting • Dubai

On-site
AED 293,793 - 440,690
Backend Engineer
Backend Engineer

AlgoDriven • Dubai

On-site
AED 350,000 - 520,000
Competitive salary
Long-term product work
Tools stipend
+1
Senior Software Engineer, AI Benchmarking
Senior Software Engineer, AI Benchmarking

Remotedxb • Dubai

On-site
AED 240,000 - 420,000
Principal Software Engineer
Principal Software Engineer

Remotedxb • Dubai

On-site
AED 360,000 - 600,000
Competitive compensation & equity
Generous total rewards
Vacation & sick days
+2
AI Software Engineer
AI Software Engineer

Janus Digital Global • Dubai

On-site
AED 360,000 - 560,000
Competitive compensation
Specialized AI domain
Real estate portfolio exposure
+3
AI Engineer (Applied)
AI Engineer (Applied)

BlackStone eIT • Dubai

On-site
AED 120,000 - 150,000
Paid Time Off
Performance Bonus
Training & Development
Full-Stack Developer AI-Powered SaaS Platform (D50 AI)
Full-Stack Developer AI-Powered SaaS Platform (D50 AI)

Division50 • Dubai

Hybrid
AED 120,000 - 180,000
Equity stake (0.5%-2%)
Remote flexibility
Direct access to founder