Turn this role into an interview — a resume and cover letter built around what this employer wants.
HumanitApp is seeking a role focused on evaluating benchmark tasks used to train and assess frontier AI models, focusing on quality and reproducibility.
You will assess tasks, patches, and harnesses while delivering rubric-based feedback to ensure rigorous evaluation standards.
This remote-friendly opportunity offers compensation ranging from $70 to $90 per hour, with workload and terms clarified during the official application process.
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity
and provide clear, rubric-based written feedback.
Basic Qualificatio...
AI evaluation Python Software engineering
This opportunity may suit professionals with relevant experience in AI evaluation, Python, Software engineering. Review the official description and requirements before applying.
The listing states $70 - $90 / hour. Confirm the final rate, workload, and payment terms during the official application process.