Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.
Mindrift is building a dataset to evaluate AI coding agents by creating realistic developer environments, tickets, docs, and conversations. The role involves crafting tasks from intermediate states of these environments and defining what counts as a solved solution.
This part-time, remote freelance opportunity pays per completed task, with rates up to $40 per hour depending on qualification tier and efficiency.
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
Why this is hard:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
Educational qualifications
Academic and/or Professional Experience
Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles
Compensation:
Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $40/hr. Because payment is per task, a faster pace raises your effective hourly rate.
Why this freelance opportunity might be a great fit for you?