Python Engineer - Freelance AI Trainer

Mindrift

Philippines

On-site

PHP 3,401,481 - 5,102,222

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focusing on testing, evaluating, and improving AI systems.

In this role, you will design tasks that reveal unsafe shortcuts, write robust tests, and assess agent solutions across realistic repositories. This is project-based work, not permanent employment, with hours around 20–25 per week and compensation up to $60 per hour equivalent.

Qualifications

  • 4–5+ years in software development.
  • Strong test design for functional and integration tests.
  • Proficiency with Python and JavaScript/TypeScript.
  • Experience with coding agents like Claude Code or Codex.
  • Familiarity with GitHub PRs and CI workflows.
  • English proficiency level at least B2+.

Responsibilities

  • Design tasks and datasets for AI coding agents.
  • Write tests that verify safe and correct agent behavior.
  • Create realistic development environments and datasets.
  • Iterate on tasks and tests based on QA feedback.

Skills

4–5+ years software development
Python
JavaScript/TypeScript
Test design (functional & integration)
Coding agents experience
GitHub PRs & CI workflows
Full-stack exposure
English proficiency (B2+)

Tools

Claude Code
GitHub Copilot CLI
Codex
GitHub CI

Job description

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.Participation isproject-based, not permanent employment.

What this opportunity involves

Frontier coding agents are already good at passing tests. We measure whether they pass themthe right way.We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:

  • Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
  • Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
  • Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this isNOT:
  • Not data labeling;
  • Not prompt engineering;
  • Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
  • Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
  • 4–5+ years in software development;
  • Core stack: Python, JavaScript/TypeScript;
  • Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
  • Familiarity with GitHub PRs and CI workflows as a user;
  • Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
  • English proficiency — B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building thetemptation— a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.

How it works
Project time expectations

For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

On this project, contributors can earn up to$60 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Python Engineer - AI Coding Agent Evaluation Freelance
Senior Python Engineer - AI Coding Agent Evaluation Freelance

Mindrift • Manila

On-site
PHP 7,630,000 - 12,716,000
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift • Philippines

On-site
Freelance Agent Evaluation Engineer
Freelance Agent Evaluation Engineer

AI Chopping Block • España

On-site
Software Engineering AI Trainer (Mexico)
Software Engineering AI Trainer (Mexico)

Anyone AI • Mexico

Hybrid
Senior Python Engineer - AI Testing & Evaluation (Contract)
Senior Python Engineer - AI Testing & Evaluation (Contract)

Mindrift • Philippines

On-site
Remote Freelance AI Evaluation & QA Engineer
Remote Freelance AI Evaluation & QA Engineer

Mindrift • Manila

On-site
PHP 1,526,000 - 2,543,000
AI Agent Evaluation Engineer — Test Architect
AI Agent Evaluation Engineer — Test Architect

Mindrift • Philippines

On-site
Senior Python Engineer - AI Evaluation Task Designer (Contract)
Senior Python Engineer - AI Evaluation Task Designer (Contract)

Mindrift • Manila

On-site
PHP 7,630,000 - 12,716,000
Senior Python Engineer — AI Task Design (Project-Based)
Senior Python Engineer — AI Task Design (Project-Based)

Mindrift • Philippines

On-site
Claims Processing Agent - Freelance AI Trainer
Claims Processing Agent - Freelance AI Trainer

Mindrift • Philippines

On-site