An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Mindrift connects specialists with project-based AI opportunities for leading tech companies. You will contribute to building realistic developer environments, crafting AI evaluation tasks, and writing robust tests.
This is a part-time, remote freelance role focused on QA automation and evaluation of AI coding agents. Ideal candidates have 5+ years in software development and a core stack including Python (FastAPI), React, Docker, Postgres, Kafka, and Redis, with strong testing experience and
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You’ll create challenging tasks and evaluation criteria within realistic simulated environments:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles.
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks - up to the equivalent of $40/hr. Because payment is per task, a faster pace raises your effective hourly rate.