Software Engineering Evaluation Specialist

Mindrift

Kuwait

On-site

KWD 8,900 - 15,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Paid contributions
Incentive payments

Job summary

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This role designs real-world coding tasks delivered as Docker environments with broken code and automated tests.

You will deliver full task packages, instructions, and reference solutions proving solvability. Participation is project-based, not permanent employment, with compensation tied to contribution and project scope.

Qualifications

  • 3+ years of production software development in Python, Go, or Node.
  • Fluent Python and pytest required.
  • Experience authoring test harnesses and fixtures.

Responsibilities

  • Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature.
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands.
  • Write an instruction and a Jira-like task description.
  • Create a reference solve and a runnable sh script.
  • Calibrate difficulty for state-of-the-art agents.
  • Iterate based on expert QA feedback.

Skills

Python
pytest
Backend development
Docker
Linux Bash

Tools

Docker

Job description

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.

Participation is project-based, not permanent employment.

About the Role

You’ll design coding tasks that challenge frontier AI coding agents.

Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome.

Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities
  • Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.
  • Write a instruction.
  • md that reads like a Jira ticket a developer would receive.
  • Write a reference solve.
  • sh proving the task is solvable.
  • Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.
  • Iterate based on feedback from expert QA reviewers.
  • Later: review other authors’ tasks as a QA reviewer.
  • Not in scope Data labeling, prompt engineering.
  • Production code to ship — you design problems and verification for AI agents.
  • Leetcode puzzles — scenarios must look like real developer work.
  • Not every candidate task ships — quality over quantity.
Requirements
  • 3+ years of production software development in one backend stack — Python, Go, Node.
  • js, Java, or Rust.
  • Depth in one stack beats breadth.
  • Python + pytest fluency — required regardless of primary stack.
  • The task harness is pytest-based even when the broken app is in another language.
  • Fixtures, parametrize, monkeypatch, timeouts, conftest.
  • py. Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work.
  • You can cite a specific time the AI was confidently wrong and how you caught it.
  • English — B2+ written.
  • Not a fit Data Science, ML, or Computer Vision engineers without backend-engineering output.
  • Manual QA testers without automation or test authoring.
  • Frontend-only, low-code / no-code, IT Support, or Business Analysts.
  • Engineers who have never written pytest from scratch.
  • Junior, intern, or assistant as the most recent role.
Preferred qualifications
  • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
  • Modern Python tooling (uv, poetry, pyproject.
  • toml). Coverage tooling (pytest-cov, coverage.
  • py, gcov, llvm-cov, kcov).
  • Fuzzing or property-based testing (Hypothesis).
  • Prior contribution to agent-evaluation benchmarks or related frameworks.
Time commitment

Onboarding: ~10 hours per first task.

Steady state: ~5 hours per task, 2–4 parallel tasks per author.

Realistic weekly load: 8–20 hours.

Higher volume available for top performers.

You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.

Compensation
  • Paid contributions, rates up to $35/hour *.
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.
  • *Rates vary based on expertise, skills assessment, location, project needs, and other factors.
  • Higher rates may be provided to highly specialized experts.
  • Lower rates may apply during onboarding or non-core project phases.
  • Payment details are shared per project.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Agent Evaluation Engineer - Remote, Freelance
AI Agent Evaluation Engineer - Remote, Freelance

Mindrift • Kuwait

On-site
KWD 13,000 - 21,000
Remote work
Freelance project paid per task
Flexible schedule
Competitive Programming Expert - Freelance AI Trainer
Competitive Programming Expert - Freelance AI Trainer

Mindrift • Kuwait

On-site
KWD 17,000 - 38,000
AI Coding Challenge Architect
AI Coding Challenge Architect

Mindrift • Kuwait

On-site
KWD 8,900 - 15,000
Paid contributions
Incentive payments
Competitive Programming Problem Architect for AI Benchmarks
Competitive Programming Problem Architect for AI Benchmarks

Mindrift • Kuwait

On-site
KWD 17,000 - 38,000
Remote Electrical Engineer - Python & AI Problem Designer
Remote Electrical Engineer - Python & AI Problem Designer

Mindrift • Kuwait

On-site
KWD 14,000 - 20,000
Fully remote
Part-time freelance
Full Stack Engineer (TypeScript / React / Node.js)
Full Stack Engineer (TypeScript / React / Node.js)

Obytes, Inc. • Kuwait

Remote
KWD 18,000 - 25,000
AI Solutions Architect (Team & Technical Lead)
AI Solutions Architect (Team & Technical Lead)

TAT IT Technolgies • Kuwait City

On-site
KWD 18,000 - 30,000
null
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics

SupportFinity™ • Kuwait City

On-site
KWD 6,000 - 10,000
Distributed work environment with in-person team sprints
Personal learning and development budget of USD 2,000 per year
Annual compensation review
+1
Entrepreneur in Residence (EIR)
Entrepreneur in Residence (EIR)

The Flex • Kuwait

On-site
KWD 24,000 - 38,000
Competitive salary
Performance-based incentives
Mentorship and strategic support
Software Engineer - Python - Ubuntu Pro client - graduate level
Software Engineer - Python - Ubuntu Pro client - graduate level

Canonical • Kuwait City

Remote
KWD 12,000 - 18,000
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Maternity and paternity leave
+1