Get more replies from employers
Send a job-specific resume in minutes.
Obsidian in San Francisco seeks a reviewer and assessor of benchmark tasks for frontier AI labs. You will evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate models, and review repository-level tasks, reference patches, and test harnesses.
You will provide rubric-based written feedback, detect leakage or reward hacking, and help improve grading integrity across the benchmark suite.
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.