Turn this role into an interview — a resume and cover letter built around what this employer wants.
Mindrift connects specialists with project-based AI opportunities for leading tech companies. You will create challenging tasks and evaluation criteria within realistic simulated environments, forming believable development histories and ensuring tasks are solvable by AI agents.
The role emphasizes writing tests that accept multiple correct solutions and requires 5+ years in software development plus strong Python/JS stack experience.
Description Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
Up to $50/hr equivalent , depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.