Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift

Portugal

Presencial

EUR 108 576 - 180 960

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focusing on testing, evaluating, and improving AI systems.

We are building a dataset to evaluate AI coding agents by designing tasks, environments, and tests, and by iterating based on QA feedback. This contract-based role is not permanent employment and emphasizes outcomes over long-term commitments.

Qualificações

  • 8+ years in software development.
  • Proficient with Python (FastAPI) and JavaScript/TypeScript (React).
  • Experience writing functional and integration tests.
  • English proficiency – B2+.

Responsabilidades

  • Build realistic developer environments – a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history.
  • Design tasks from intermediate states of these environments – craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions – accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback – review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Conhecimentos

Software development
Python (FastAPI)
JavaScript/TypeScript (React)
Testing (functional & integration)
English (B2+)

Ferramentas

Docker
PostgreSQL
Kafka
Redis

Descrição da oferta de emprego

Mindrift Connects Specialists with Project‑Based AI Opportunities

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

What this opportunity involves

We're building a dataset to evaluate AI coding agents – how well a model handles real‑world developer tasks.

Responsibilities
  • Build realistic developer environments – a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments – craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions – accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback – review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch – the agent writes most of the code; you guide and evaluate
Qualifications
  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency – B2+
Challenges

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non‑trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions – writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How it works

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

Effort estimate

Task Estimates

Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

Up to $150/hr equivalent, depending on level and pace. Tasks are estimated at ~30 hours each; you set your own schedule.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Python Engineer - Freelance AI Trainer
Python Engineer - Freelance AI Trainer

Mindrift • Portugal

Presencial
Freelance AI Evaluation Engineer (Python/Full-Stack)
Freelance AI Evaluation Engineer (Python/Full-Stack)

Mindrift • Portugal

Presencial
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Lisboa

Teletrabalho
Flexible remote work
Competitive rates
Experience in AI projects
Senior Python Engineer — AI Task Designer & Evaluator
Senior Python Engineer — AI Task Designer & Evaluator

Mindrift • Portugal

Presencial
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Lisboa

Teletrabalho
Competitive pay up to $30/hour
Remote work flexibility
Gain valuable AI project experience
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Portugal

Teletrabalho
Flexible working hours
Competitive hourly rates
Experience on advanced AI projects
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Lisboa

Teletrabalho
Flexible project hours
Competitive hourly rates
Gain experience in advanced AI projects
Freelance Financial Analyst - AI Trainer
Freelance Financial Analyst - AI Trainer

Mindrift • Lisboa

Teletrabalho
Flexible work schedule
Competitive pay up to $34/hour
Experience on advanced AI projects
Freelance AI Trainer - Civil Engineering & Python
Freelance AI Trainer - Civil Engineering & Python

Mindrift • Lisboa

Teletrabalho
Flexible work hours
Competitive hourly rate
Freelance Finance Expert - AI Trainer
Freelance Finance Expert - AI Trainer

Mindrift • Portugal

Teletrabalho
Competitive pay up to $33/hour
Remote, flexible work schedule
Involvement in advanced AI projects