Freelance Agent Evaluation Engineer

AI Chopping Block

Buenos Aires

Híbrido

ARS 31.166.000 - 62.331.000

A tiempo parcial

Hace 12 días
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Flexible hours
Remote freelance project

Descripción de la vacante

Mindrift is seeking a seasoned software professional to design and evaluate AI developer tasks on a project-based basis. You will craft realistic environments, specify what counts as solved, and write robust tests to measure agent performance.

This part-time, remote freelance role emphasizes deep understanding of modeling failures and creating fair evaluation criteria across Python, React, and containerized tech stacks.

Formación

  • 5+ years in software development.
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis.
  • Experience writing tests (functional, integration).
  • English proficiency - B2+.

Responsabilidades

  • Create challenging tasks and evaluation criteria within realistic simulated environments.
  • Write tests that verify agent solutions and accept multiple valid approaches.
  • Iterate on tasks and tests based on QA feedback to ensure fairness.
  • Review agent solutions, analyze failures, and refine prompts and tests.

Conocimientos

Software development
QA testing
English B2+
Task design

Educación

Master's Degree in Computer Science / Software Engineering
Bachelor’s degree (optional with 5 years experience)

Herramientas

Python (FastAPI)
JavaScript/TypeScript (React)
Docker
Postgres
Kafka
Redis

Descripción del empleo

Mindrift connects specialists with project‑based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.Participation isproject-based, not permanent employment.

We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

You'll create challenging tasks and evaluation criteria within realistic simulated environments:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What we look for:
  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+
Why this is hard:

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

Requirements and benefits
Educational qualifications
  • A Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields.
  • Bachelor’s degree is accepted if only candidate has 5 years of experience in the field.
Academic and/or Professional Experience

Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles

How it works

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

Compensation:

Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $30/hr. Because payment is per task, a faster pace raises your effective hourly rate.

Why this freelance opportunity might be a great fit for you?
  • Take part in a part-time, remote, freelance project that fits around your primary professional or academic commitments.
  • Work on advanced AI projects and gain valuable experience that enhances your portfolio. - Influence how future AI models understand and communicate in your field of expertise.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Freelance AI Evaluation Engineer
Remote Freelance AI Evaluation Engineer

AI Chopping Block • Buenos Aires

Híbrido
ARS 31.166.000 - 62.331.000
Flexible hours
Remote freelance project
AI Agent Evaluation Engineer - Project-Based
AI Agent Evaluation Engineer - Project-Based

Mindrift • Argentina

Presencial
Senior AI Engineer
Senior AI Engineer

Athenaworks • Argentina

Presencial
ARS 176.967.000 - 265.452.000
Payment in USD
Flexible work schedule
Non-working pay days
AI Engineer
AI Engineer

Aqusag-Technologies • Argentina

A distancia
USD 120.000 - 170.000
Software Engineer (AI Training)
Software Engineer (AI Training)

Alignerr • Ciudad de Mendoza

A distancia
ARS 83.202.000 - 166.404.000
Fully remote
Flexible hours
Freelance autonomy
+1
Senior AI Agent Engineer
Senior AI Agent Engineer

Wizeline • Argentina

Presencial
ARS 3.000.000 - 5.400.000
A High-Impact Environment
Commitment to Professional Development
Flexible and Collaborative Culture
+3
Senior Python AI Engineer
Senior Python AI Engineer

Proxify • Argentina

A distancia
ARS 123.440.000 - 145.225.000
Flexible withdrawal options
Consistent 8-hour working days
Up to 24 flex days off per year
AI Automation Specialist (Part time)
AI Automation Specialist (Part time)

Scalia • Ciudad de Mendoza

A distancia
Confidential
Compensation paid in USD
Flexible and remote-first environment
Opportunity to work on cutting-edge projects
+1
Software Engineer (AI Training)
Software Engineer (AI Training)

Alignerr • La Plata

A distancia
ARS 125.058.000 - 250.117.000
Principal AI Engineer
Principal AI Engineer

Creative Chaos • Argentina

Presencial
ARS 1.800.000 - 3.000.000