Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Mindrift is assembling a dataset and evaluation framework to test AI coding agents. You will construct realistic developer environments, including codebases, infrastructure, and contextual artifacts such as tickets and docs, to simulate a believable history.
You will design tasks from intermediate states, define solvable criteria for AI agents, and craft tests that tolerate multiple valid approaches while rejecting incorrect ones.
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at :20 hours each; you set your own schedule.