Get more replies from employers
Send a job-specific resume in minutes.
Mercor seeks a professional to evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate frontier AI models. You will assess repository-level tasks, reference patches, test harnesses, and grading integrity.
You will deliver clear rubric-based written feedback, outlining strengths, gaps, and concrete improvements to ensure reliable benchmarking and reproducibility across teams working with open-source tools.
Mercor seeks a professional to evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate frontier AI models. You will assess repository-level tasks, reference patches, test harnesses, and grading integrity.
You will deliver clear rubric-based written feedback, outlining strengths, gaps, and concrete improvements to ensure reliable benchmarking and reproducibility across teams working with open-source tools.