Get more replies from employers
Send a job-specific resume in minutes.
Mercor, in partnership with a leading AI research organization, seeks a seasoned data scientist to define what excellent work looks like by designing task-specific grading criteria and scoring sample results with rigorous, written justifications.
You will evaluate AI-generated and human work, ensure scores are reproducible and defensible, and incorporate structured feedback from senior reviewers to iterate on standards for real-world data science deliverables.
Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.
Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)
Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
Apply consistent, evidence-based judgment so that scores are reproducible and defensible
Incorporate structured feedback from senior reviewers and iterate quickly on your work
5+ years of professional data science experience in industry
Background in business operations, product, or growth data science at top-tier technology companies
Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders
Exceptionally strong written communication
Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
Prior experience with AI training, evaluation, or human-data projects is a strong plus