Software Engineering Evaluation Specialist

Mindrift

Glasgow

On-site

GBP 21,000 - 35,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift in Glasgow, Scotland is seeking a Software Engineering Evaluation Specialist to design AI-task challenges. You will craft self-contained Docker environments containing broken code, tests, instructions, and a reference solution.

You’ll ensure reproducibility and deterministic verification with pytest, while producing clear Jira-style tickets and a solvable reference script. The role emphasizes project-based work rather than permanent employment, with on-demand contributions and

Qualifications

  • 3+ years of production software development in a backend stack.
  • Fluency in Python and pytest required regardless of primary language.
  • Experience with Docker-based task environments and Linux debugging.

Responsibilities

  • Design reproducible Docker environments with pinned dependencies.
  • Write deterministic pytest tests that verify outcomes without leaking fixes.
  • Create instruction.md styled as a Jira ticket for developers.
  • Provide a reference solve.sh proving task solvability.
  • Calibrate task difficulty for 20–60% success by current AI agents.
  • Iterate based on QA reviewer feedback; later act as QA reviewer for others.

Skills

Python
pytest
Docker
Linux
Bash
AI coding agent experience

Tools

Docker
pytest
CI/CD

Job description

Software Engineering Evaluation Specialist Mindrift•Glasgow, Scotland, GB

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

About the Role

You’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.
  • Write an instruction.md that reads like a Jira ticket a developer would receive.
  • Write a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.
  • Iterate based on feedback from expert QA reviewers.
  • Later: review other authors’ tasks as a QA reviewer.
Not in scope
  • Production code to ship — you design problems and verification for AI agents.
  • Not every candidate task ships — quality over quantity.
Requirements
  • 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
  • Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
  • Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
Not a fit
  • Data Science, ML, or Computer Vision engineers without backend-engineering output.
  • Manual QA testers without automation or test authoring.
  • Frontend-only, low-code / no-code, IT Support, or Business Analysts.
  • Engineers who have never written pytest from scratch.
  • Junior, intern, or assistant as the most recent role.
Preferred qualifications
  • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
  • Modern Python tooling (uv, poetry, pyproject.toml).
  • Fuzzing or property-based testing (Hypothesis).
  • Prior contribution to agent-evaluation benchmarks or related frameworks.
Process
Time commitment
  • Onboarding: ~10 hours per first task.
  • Realistic weekly load: 8–20 hours. Higher volume available for top performers.
  • You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
Compensation
  • Paid contributions, rates up to $35/hour*.
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.

*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • United Kingdom

Remote
Task-based compensation
Fully remote work
Flexible participation hours
+1
Evaluation Scenario Writer - AI Agent Testing Specialist
Evaluation Scenario Writer - AI Agent Testing Specialist

Mindrift • United Kingdom

Remote
Competitive hourly rates
Flexible work schedule
Experience in an advanced AI project
Freelance Electrical Engineering Specialist (Python) - AI Trainer
Freelance Electrical Engineering Specialist (Python) - AI Trainer

Mindrift • United Kingdom

Remote
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • United Kingdom

Remote
Competitive hourly rates up to $50
Flexible working schedule
Work on cutting-edge AI projects
+1
Freelance Machine Learning Engineer
Freelance Machine Learning Engineer

Mindrift • Manchester

On-site
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity
Private healthcare
+2
AI Benchmark Engineer: Backend Task Design & Evaluation
AI Benchmark Engineer: Backend Task Design & Evaluation

Mindrift • Glasgow

On-site
GBP 21,000 - 35,000
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • United Kingdom

Remote
Competitive hourly rates up to $50
Flexible part-time hours
Experience with advanced AI projects
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • United Kingdom

Remote
GBP 31,000 - 51,000
Competitive pay up to $50/hour
Flexible, part-time remote work
Experience in advanced AI projects
Electrical Engineer with Python Experience - Freelance AI Trainer
Electrical Engineer with Python Experience - Freelance AI Trainer

Mindrift • Helensburgh

Remote
Flexible schedule
Remote work
Competitive pay