Get more replies from employers
Send a job-specific resume in minutes.
Mercor is seeking experienced data scientists to define grading criteria for real-world data science deliverables and to score AI-generated and human work samples with rigorous, written justifications. The role focuses on designing task-specific evaluation criteria and ensuring that scores are reproducible and defensible.
You will apply evidence-based judgment, incorporate feedback from senior reviewers, and iteratively refine scoring strategies while communicating findings to executive
Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real‑world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task‑specific grading criteria and scoring completed work samples with rigorous, well‑reasoned written justifications.