Turn this role into an interview — a resume and cover letter built around what this employer wants.
Obsidian in San Francisco seeks a reviewer and assessor of benchmark tasks for frontier AI labs. You will evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate models, and review repository-level tasks, reference patches, and test harnesses.
You will provide rubric-based written feedback, detect leakage or reward hacking, and help improve grading integrity across the benchmark suite.
Obsidian in San Francisco seeks a reviewer and assessor of benchmark tasks for frontier AI labs. You will evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate models, and review repository-level tasks, reference patches, and test harnesses.
You will provide rubric-based written feedback, detect leakage or reward hacking, and help improve grading integrity across the benchmark suite.