Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Obsidian is seeking a senior professional to design and run evaluation tasks that test AI models against expert-level work in materials science or electrical/mechanical engineering. You will craft prompts, data rooms, and grading schemes to produce objective measurements of model performance.
The role involves shaping undefined work streams, collaborating with domain experts, and iterating on criteria to clearly distinguish good vs. suboptimal model reasoning.
About the role
We build materials-science and engineering tasks that test how well an AI model does real expert work. You write a realistic task — a prompt, a data room, and a way to grade it — run it against the model, and tighten it until the model can no longer reason through it cleanly.
What you'll do
Who we're looking for
Support
Daily onboarding sessions and multiple office hours every day.