Remote Data Science & AI Evaluation Benchmark Designer

Weekday AI

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Weekday AI is seeking experienced Data Scientists and Quantitative Analysts to design realistic benchmark tasks for frontier AI models. You will clean data, perform statistical analyses, and interpret results to inform evaluations, working closely with AI researchers in a fully remote, full-time role.

The position requires 1+ year in data-intensive roles, strong Python skills, and experience with Jupyter Notebooks or Colab. Expect ~35 hours weekly, with weekly payments for completed work.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in Data Science, Statistics, Mathematics, Economics, Operations Research, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in research, research engineering, quantitative analysis, data science, or another data-intensive analytical role.
  • Strong hands-on experience with data cleaning, exploratory data analysis, statistical testing, correlation analysis, predictive modeling, and interpretation of analytical results.
  • Proficiency with Jupyter Notebooks or Google Colab for building reproducible analytical workflows.
  • Strong programming skills in Python, including pandas, NumPy, SciPy, scikit-learn, or similar analytical frameworks.
  • Working knowledge of Git and collaborative software development practices.
  • Excellent written communication skills with the ability to present analytical findings clearly to both technical and non-technical audiences.
  • Experience with AI evaluation, benchmark development, AI model assessment, or task authoring is preferred.
  • Exceptional analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended problems independently.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Responsibilities

  • Design realistic data analysis challenges inspired by real-world research and analytical workflows, including data preparation, statistical modeling, hypothesis testing, and comparative analysis.
  • Develop reproducible reference analyses using Jupyter Notebooks or Google Colab, documenting methodologies and findings with clarity.
  • Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations.
  • Evaluate AI-generated analyses for correctness, statistical validity, reasoning quality, and interpretation accuracy.
  • Identify analytical errors, flawed assumptions, and reasoning gaps that experienced data professionals would immediately recognize.
  • Collaborate with AI researchers and fellow subject matter experts to improve benchmark quality, consistency, and analytical rigor.

Skills

Data Science
Statistics
Mathematics
Economics
Operations Research
Analytical thinking

Education

Master's degree, PhD or equivalent in quantitative STEM

Tools

Jupyter Notebooks
Google Colab
Git
Python
pandas
NumPy
SciPy
scikit-learn

Job description

Weekday AI is seeking experienced Data Scientists and Quantitative Analysts to design realistic benchmark tasks for frontier AI models. You will clean data, perform statistical analyses, and interpret results to inform evaluations, working closely with AI researchers in a fully remote, full-time role.

The position requires 1+ year in data-intensive roles, strong Python skills, and experience with Jupyter Notebooks or Colab. Expect ~35 hours weekly, with weekly payments for completed work.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Data Science & AI Benchmark Analyst
Remote Data Science & AI Benchmark Analyst

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
AI Benchmark Engineer (Python) — Remote
AI Benchmark Engineer (Python) — Remote

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Benchmark Architect (Remote) | Python
AI Benchmark Architect (Remote) | Python

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
Remote AI Benchmark Researcher - Computational STEM
Remote AI Benchmark Researcher - Computational STEM

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Evaluation Architect - Data Science Expert (Remote)
AI Evaluation Architect - Data Science Expert (Remote)

Weekday AI (YC W21) • United States

On-site
USD 165,312 - 234,192
Fully remote
Weekly payments