Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.
Mercor seeks a research-driven professional to design frontier AI evaluation problems in drug research. You will craft realistic challenges, assemble a source-grounded data room, and author a rigorous grader to assess submissions on substance and reasoning.
The role emphasizes defense of conclusions in complex, evidence-driven settings. Minimum commitment is 20 hours per week, up to 40, with browser-based tooling and Claude Max access provided.
This is an evaluation framework for frontier AI models in drug research and development. Domain experts write realistic research problems wrapped around messy evidence worlds. The model works offline in a sandbox with open-source scientific tooling, receiving incomplete, indirect and sometimes misleading data, and must reach a defensible conclusion. The answer is present in the evidence but never stated.
Your job is to build a research problem so realistic and well-constructed that the model has to genuinely reason to solve it — it can't pattern-match or look the answer up.
Three things per task: the prompt, the data room, and the grading document.
Design the challenge — identify the hidden basin and the forcing chain that makes one conclusion inevitable.
Build the world — assemble a data room from real, source‑grounded data. Every decisive value traces to a real source.
Write the grader — a grading document that scores a submitted memo on substance, not style or method.
Validate it — run verification and calibration before submission.
Revise it — you own the rework if a reviewer requests changes.
You are the final authority on every scientific and difficulty question in your task.
Working expertise in mechanistic enzymology, structure‑guided peptide or antibody design, fragment‑based discovery, or structure‑based medicinal chemistry — PhD or equivalent industry track record.
You can pose a problem whose answer you can defend rigorously, but which a strong solver cannot shortcut.
Comfortable judging reasoning quality, not just final answers.
Minimum 20 hours per week, up to 40.
Work happens in a browser‑based studio plus Claude Code. A Claude Max subscription is required and is fully reimbursed.