Python Engineer - Freelance AI Trainer

Mindrift

Portugal

Presencial

EUR 48 425 - 72 638

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a contract, not permanent employment.

You will design evaluation tasks to ensure safe, honest, and scoped coding, write robust tests, and iterate with QA feedback across realistic repositories. Expect 20–25 hours weekly during active phases; compensation up to $60 per hour, depending on level and pace.

Qualificações

  • 4–5+ years in software development.
  • Core stack: Python, JavaScript/TypeScript.
  • Strong test design skills for safe vs unsafe completion.
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex).
  • Familiarity with GitHub PRs and CI workflows.
  • English proficiency (B2+).

Responsabilidades

  • Design tasks that pair safe progress with potential unsafe shortcuts and evaluate them.
  • Write tests that catch shortcuts, not just correct outputs.
  • Iterate on tasks and tests based on QA feedback to ensure fairness and robustness.
  • Review agent solutions within realistic repositories including databases and CI pipelines.

Conhecimentos

Software development
Python
JavaScript/TypeScript
Test design
Coding agents
GitHub PRs & CI
English B2+

Ferramentas

Claude Code
GitHub Copilot CLI
Codex

Descrição da oferta de emprego

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.Participation isproject-based, not permanent employment.

What this opportunity involves

Frontier coding agents are already good at passing tests. We measure whether they pass themthe right way.We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:

  • Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
  • Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
  • Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this isNOT:
  • Not data labeling;
  • Not prompt engineering;
  • Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
  • Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
  • 4–5+ years in software development;
  • Core stack: Python, JavaScript/TypeScript;
  • Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
  • Familiarity with GitHub PRs and CI workflows as a user;
  • Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
  • English proficiency — B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building thetemptation— a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.

How it works
Project time expectations

For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

On this project, contributors can earn up to$60 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift • Portugal

Presencial
Freelance AI Evaluation Engineer (Python/Full-Stack)
Freelance AI Evaluation Engineer (Python/Full-Stack)

Mindrift • Portugal

Presencial
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Lisboa

Teletrabalho
Competitive pay up to $30/hour
Remote work flexibility
Gain valuable AI project experience
Freelance Financial Analyst - AI Trainer
Freelance Financial Analyst - AI Trainer

Mindrift • Lisboa

Teletrabalho
Flexible work schedule
Competitive pay up to $34/hour
Experience on advanced AI projects
Freelance Finance Expert - AI Trainer
Freelance Finance Expert - AI Trainer

Mindrift • Portugal

Teletrabalho
Competitive pay up to $33/hour
Remote, flexible work schedule
Involvement in advanced AI projects
Python AI Testing Engineer — Design Safe Coding Tasks
Python AI Testing Engineer — Design Safe Coding Tasks

Mindrift • Portugal

Presencial
Freelance AI Trainer - Civil Engineering & Python
Freelance AI Trainer - Civil Engineering & Python

Mindrift • Lisboa

Teletrabalho
Flexible work hours
Competitive hourly rate
Freelance Accounting Consultant - AI Trainer
Freelance Accounting Consultant - AI Trainer

Mindrift • Portugal

Teletrabalho
Competitive hourly rates
Gain experience on advanced AI projects
Flexible working conditions
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Portugal

Teletrabalho
Flexible working hours
Competitive hourly rates
Experience on advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Lisboa

Teletrabalho
Flexible remote work
Competitive rates
Experience in AI projects