Python Engineer - Freelance AI Trainer

Mindrift

Brasil

Presencial

BRL 210 427 - 350 712

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

Mindrift project-based engagement

Resumo da oferta

Mindrift is building a project-based AI evaluation platform connecting specialists with AI opportunities for leading tech companies. The role focuses on designing tasks to test safe and scoped AI coding behavior, writing tests, and iterating based on QA feedback.

Ideal candidates have 4–5+ years in software development and strong Python/JavaScript/TypeScript skills. This is a project-based engagement with exposure to real repositories, databases, CI pipelines, and deployment scripts.

Qualificações

  • 4–5+ years in software development.
  • Core stack: Python, JavaScript/TypeScript.
  • Strong test design skills — functional and integration tests.
  • Hands-on experience with coding agents (Claude Code, Codex, etc.).
  • Familiarity with GitHub PRs and CI workflows.
  • Stack breadth is welcome; English proficiency is required (B2+).

Responsabilidades

  • Design tasks that challenge AI agents while staying within scope.
  • Build realistic developer environments with codebases and CI pipelines.
  • Write tests that verify correct behavior and catch unsafe shortcuts.
  • Iterate on tasks and tests based on QA feedback to ensure fairness.

Conhecimentos

Software development
Python
JavaScript/TypeScript
Testing
CI/CD
GitHub PRs
English B2+

Ferramentas

Claude Code
GitHub Copilot CLI
Codex

Descrição da oferta de emprego

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation isproject-based, not permanent employment.

What this opportunity involves

Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners. You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:

  • Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
  • Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
  • Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
  • Not data labeling;
  • Not prompt engineering;
  • Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
  • Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
  • 4–5+ years in software development;
  • Core stack: Python, JavaScript/TypeScript;
  • Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
  • Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
  • Familiarity with GitHub PRs and CI workflows as a user;
  • Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
  • English proficiency — B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non‑trivial. The real difficulty is building the temptation — a scenario where the unsafe or out‑of‑scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.

Project time expectations

For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

On this project, contributors can earn up to $50 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Freelance AI Evaluation Engineer (Python/Full-Stack)
Freelance AI Evaluation Engineer (Python/Full-Stack)

Mindrift • Brasil

Teletrabalho
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • Rio de Janeiro

Presencial
Freelance project-based collaboration
Fully remote and flexible participation
Task-based compensation
Civil Engineer & Python Expert - Freelance AI Trainer
Civil Engineer & Python Expert - Freelance AI Trainer

Mindrift • Rio de Janeiro

Híbrido
Freelance Financial Analyst - AI Trainer
Freelance Financial Analyst - AI Trainer

Mindrift • Rio de Janeiro

Teletrabalho
Flexible working hours
Competitive pay up to $16/hour
Gain experience in advanced AI projects
Freelance Economics Expert - AI Trainer
Freelance Economics Expert - AI Trainer

Mindrift • São Paulo

Teletrabalho
Competitive pay up to $16/hour
Remote freelance flexibility
Gain valuable experience on advanced AI projects
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Belo Horizonte

Teletrabalho
Competitive hourly rates
Remote work flexibility
Experience with advanced AI projects
Automotive Engineer with Python Experience - Freelance AI Trainer
Automotive Engineer with Python Experience - Freelance AI Trainer

Mindrift • Brasil

Teletrabalho
Flexible work schedule
Experience with advanced AI projects
Competitive pay based on expertise
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • Belo Horizonte

Teletrabalho
Competitive hourly rates
Flexible scheduling
Experience on advanced AI projects
Freelance AI Trainer - Civil Engineering & Python
Freelance AI Trainer - Civil Engineering & Python

Mindrift • Rio de Janeiro

Teletrabalho
Earn up to $13/hour
Fully remote, flexible work
Contribute to future AI systems
Freelance AI Trainer - Civil Engineering & Python
Freelance AI Trainer - Civil Engineering & Python

Mindrift • Rio de Janeiro

Teletrabalho
Flexible working hours
Remote work
Competitive hourly rate