Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.

You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.

Qualifications

  • Master's degree or PhD in a STEM field or equivalent research experience.
  • At least 1 year of experience in an active research role.
  • Experience with Python, data analysis, simulation, ML, or scientific computing.
  • Strong understanding of experimental design and statistical analysis.
  • Familiarity with Git, IDEs, and notebook platforms (Jupyter or Colab).

Responsibilities

  • Design complex benchmark tasks informed by real-world scientific workflows.
  • Develop reference solutions using Python, notebooks, and computational tools.
  • Define evaluation standards to distinguish sound reasoning from flawed conclusions.
  • Review AI-generated solutions and identify methodological weaknesses.
  • Collaborate with researchers to improve benchmark quality and rigor.
  • Refine evaluation methodologies for advanced AI systems.

Skills

Python
Data analysis
Machine learning
Statistical analysis
Scientific computing

Education

Master's degree
PhD

Tools

Git
Jupyter
Google Colab

Job description

Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.

You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Data Science & AI Benchmark Analyst
Remote Data Science & AI Benchmark Analyst

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Remote AI Evaluation Architect - Tech Docs & Code
Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote Bio AI Evaluation Scientist (PhD)
Remote Bio AI Evaluation Scientist (PhD)

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote AI Evaluation Specialist
Remote AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
Remote QA/Test Engineer — AI Benchmark Validation
Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
Remote AI Model Evaluator & Benchmark Scientist
Remote AI Model Evaluator & Benchmark Scientist

OpenTrain AI, Inc. • United States

Remote
USD 34,000 - 55,000
Remote AI Consulting Strategist: Prompt Design & Benchmarking
Remote AI Consulting Strategist: Prompt Design & Benchmarking

Weekday 1 • United States

Remote
USD 110,000 - 165,000
Equity options
Weekly performance incentives (hourly)
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1