An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Obsidian seeks an experienced evaluator to assess AI-assisted software-development traces used to train frontier AI models. You will judge correctness, workflow soundness, and reasoning across end-to-end coding sessions produced with AI-assisted developer tools.
Responsibilities include rubric-based feedback and assessing trajectories for best practices, with a focus on maintaining high-quality developer tooling and workflows.
Evaluate the quality and correctness of AI-assisted software-development traces used to train and evaluate a frontier AI lab's models. You'll assess end-to-end coding sessions produced with AI-assisted developer tools — judging correctness, workflow soundness, and reasoning — and provide clear, rubric-based written feedback.