Turn this role into an interview — a resume and cover letter built around what this employer wants.
OpenTrain is seeking a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.
The role requires at least three years of software engineering experience, open-source contributions, and the ability to audit patches, test runners, and grading systems.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a profile that reflects their experience, and apply in minutes.
As an OpenTrain contractor, you can develop a durable portfolio of AI training work while finding opportunities that match your technical background. Creating an OpenTrain account is free.
AI training is the human side of building artificial intelligence. Technical contributors review code, assess model outputs, and evaluate benchmark tasks so AI systems can become more accurate, reliable, and useful.
This work puts experienced software professionals close to the development of cutting-edge AI systems. Remote projects can offer flexible schedules and the opportunity to apply practical engineering judgment to advanced model evaluation.
OpenTrain is recruiting a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.
The role combines practical software-engineering judgment with careful evaluation of whether benchmark tasks are correct, reproducible, and resistant to answer leakage or reward hacking. You will provide concise feedback grounded in defined evaluation criteria.
You will review software-engineering tasks at the repository level and assess whether they accurately represent the intended work. Your findings will help identify benchmark weaknesses that could distort AI model evaluation.
This role requires at least three years of professional software-engineering experience, along with meaningful open-source contribution or maintainer experience. The listing is marked entry level, but candidates must meet the stated professional experience and technical requirements.
This opportunity is suited to software engineers who can move comfortably between source code, test infrastructure, containerized execution, and evaluation criteria. It may be especially relevant to open-source contributors and maintainers who understand how repository-level changes should be tested and reviewed.
Strong candidates will be able to explain technical findings clearly, distinguish legitimate task difficulty from benchmark defects, and recognize when evaluation setups create opportunities for leakage or reward hacking.