Eine maßgeschneiderte Bewerbung für diese Stelle — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.
Obsidian in Berlin seeks a science-driven evaluator to design frontier AI problems in drug research. The role builds realistic data rooms and solvable yet non-patterned prompts, requiring deep domain expertise and rigorous defense of conclusions.
You will deliver prompts, data rooms, and grading documents, collaborate with domain experts, and validate work through calibration. A PhD or equivalent experience in enzymology, medicinal chemistry or structure-based design is essential.
This is an evaluation framework for frontier AI models in drug research and development. Domain experts write realistic research problems wrapped around messy evidence worlds. The model works offline in a sandbox with open-source scientific tooling, receiving incomplete, indirect and sometimes misleading data, and must reach a defensible conclusion. The answer is present in the evidence but never stated.
Your job is to build a research problem so realistic and well-constructed that the model has to genuinely reason to solve it - it can't pattern-match or look the answer up.
Three things per task: the prompt, the data room, and the grading document.
You are the final authority on every scientific and difficulty question in your task.
Minimum 20 hours per week, up to 40.
Work happens in a browser-based studio plus Claude Code. A Claude Max subscription is required and is fully reimbursed.