Senior Python Engineer - AI Coding Agent Evaluation Freelance

Mindrift

Chennai District

On-site

INR 7,910,000 - 13,183,000

Part time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focusing on testing, evaluating, and improving AI systems. You will build realistic developer environments, design tasks, and write tests for AI agents in a project-based setting.

Applicants should have 8+ years in software development, expertise in Python (FastAPI), JS/TS (React), Docker, Postgres, Kafka, Redis, and experience writing functional and integration tests, with English proficiency at B2+.

Qualifications

  • ,

Responsibilities

  • Build realistic developer environments – a virtual company with codebase, infrastructure, and context that forms a believable development history.
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Skills

8+ years
Python FastAPI
JavaScript TypeScript React
Docker
Postgres
Kafka
Redis
Testing (functional, integration)
English (B2+)

Tools

Docker
Postgres
Kafka
Redis

Job description

Job Description:


Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.What this opportunity involves:
Were building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.


Youll create challenging tasks and evaluation criteria within realistic simulated environments:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

What this is NOT:

  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

What we look for:

  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Why this is hard:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How it works
  • Apply
  • Pass qualification(s)
  • Join a project
  • Complete tasks
  • Get paidEffort estimate

Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation:
Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $100/hr. Because payment is per task, a faster pace raises your effective hourly rate.

Requirements:

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)
Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift • India

On-site
Freelance Agent Evaluation Engineer
Freelance Agent Evaluation Engineer

Mindrift • Chennai District

On-site
INR 3,258,278,167 - 5,213,245,067
Freelance Agent Evaluation Engineer
Freelance Agent Evaluation Engineer

Mindrift • Mumbai

Remote
Software Engineering Evaluation Specialist
Software Engineering Evaluation Specialist

Mindrift • Chennai District

On-site
INR 2,768,000 - 4,614,000
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • India

On-site
INR 2,886,000 - 3,936,000
Evaluation Scenario Writer - AI Agent Testing Specialist
Evaluation Scenario Writer - AI Agent Testing Specialist

Mindrift • Hyderabad

Remote
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • Mumbai

On-site
Freelance project-based collaboration
Flexible participation hours
Opportunity to work on AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Maharashtra

Remote
Flexible schedule
Competitive pay up to $12/hour
Gain experience in advanced AI projects
Freelance Agent Evaluation Analyst
Freelance Agent Evaluation Analyst

Mindrift • New Delhi

Remote
Flexible work hours
Competitive pay up to $12/hour
Experience on advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Ahmedabad District

Remote
Flexible project schedule
Competitive hourly rates
Experience in advanced AI projects