An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Explore more jobs seeks a software-testing specialist to help build evaluation tasks for AI coding agents. You will design realistic developer environments and write tests that differentiate safe from unsafe completion.
This role focuses on ensuring tests cover breadth of repos, CI workflows and codebases, with 4-5+ years in software development and hands-on agent experience.
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents - not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.
You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation - a scenario where the unsafe or out-of-scope path is the path of least resistance - and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
On this project, contributors can earn up to $75 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.