Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.
Mindrift is seeking a seasoned software professional to design and evaluate AI developer tasks on a project-based basis. You will craft realistic environments, specify what counts as solved, and write robust tests to measure agent performance.
This part-time, remote freelance role emphasizes deep understanding of modeling failures and creating fair evaluation criteria across Python, React, and containerized tech stacks.
Mindrift connects specialists with project‑based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.Participation isproject-based, not permanent employment.
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $30/hr. Because payment is per task, a faster pace raises your effective hourly rate.