Python Engineer - Freelance AI Trainer

Mindrift

Sweden

On-site

SEK 803,265 - 1,004,081

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mindrift connects specialists with project-based AI opportunities for leading tech companies. You will design tasks that test safe, honest behavior in AI coding agents, write robust tests, and iterate based on QA feedback to ensure fair evaluation across diverse solutions.

As part of a hands-on engineering role, you will build realistic test environments, review agent submissions, and refine scenarios to maintain high-quality benchmarks for safety and performance.

Qualifications

  • 4–5+ years in software development.
  • Core stack: Python, JavaScript/TypeScript.
  • Strong test design skills for safe vs unsafe evaluation.
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex).
  • Familiarity with GitHub PRs and CI workflows.
  • English proficiency at least B2+.

Responsibilities

  • Design tasks that pair a benign goal with unsafe shortcuts and write tests to catch it.
  • Review agent solutions, analyze failures, and refine tasks for fair evaluation.
  • Create realistic developer environments with codebases, infra, tickets, and docs.

Skills

Software development
Python
JavaScript/TypeScript
Test design
Coding agents
GitHub PRs & CI
English B2+

Tools

Claude Code
GitHub Copilot CLI
Codex

Job description

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.Participation isproject-based, not permanent employment.

What this opportunity involves

Frontier coding agents are already good at passing tests. We measure whether they pass themthe right way.We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:

  • Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
  • Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
  • Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
  • Not data labeling;
  • Not prompt engineering;
  • Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
  • Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
  • 4–5+ years in software development;
  • Core stack: Python, JavaScript/TypeScript;
  • Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
  • Familiarity with GitHub PRs and CI workflows as a user;
  • Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
  • English proficiency — B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building thetemptation— a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.

How it works
Project time expectations

For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

On this project, contributors can earn up to$75 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Python Engineer for AI Testing & Task Design
Python Engineer for AI Testing & Task Design

Mindrift • Sweden

On-site
Senior Python Engineer - AI Task Designer & Evaluator
Senior Python Engineer - AI Task Designer & Evaluator

Mindrift • Sweden

On-site
DevOps Engineer - AI Model Evaluator
DevOps Engineer - AI Model Evaluator

Mercor • Stockholms kommun

On-site
SEK 4,723,000 - 5,773,000
Full-Stack Developer - AI Trainer (Sweden)
Full-Stack Developer - AI Trainer (Sweden)

Anyone AI • Stockholms kommun

On-site
Remote Full-Stack Engineer for AI Evaluation Projects
Remote Full-Stack Engineer for AI Evaluation Projects

Anyone AI • Stockholms kommun

On-site
Senior Python AI Engineer
Senior Python AI Engineer

Klarna • Stockholms kommun

On-site
SEK 500,000 - 700,000
Applied AI Engineer - Agents & Local Models
Applied AI Engineer - Agents & Local Models

Playground Group • Stockholms kommun

On-site
SEK 600,000 - 800,000
Remote AI Task Designer for Software Engineers
Remote AI Task Designer for Software Engineers

Anyone AI • Stockholms kommun

On-site
Remote Strategy Consultant for AI Training
Remote Strategy Consultant for AI Training

Mindrift • Sweden

Remote
Remote work
Project-based engagement
Competitive hourly pay
AI DevOps Evaluator for Frontier Code Agents
AI DevOps Evaluator for Frontier Code Agents

Mercor • Stockholms kommun

On-site
SEK 4,723,000 - 5,773,000