Remote AI Benchmark Researcher - Computational STEM

Weekday AI

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday AI is hiring a researcher for a fully remote, full-time role in the United States. You will design complex benchmark tasks, develop reference solutions in Python, and help uncover methodological flaws in state-of-the-art AI models.

The position emphasizes rigorous experimental design, data analysis, and collaboration with AI researchers. Qualifications include a STEM terminal degree or equivalent experience, 1+ year in research, strong Python/data analysis skills, and familiarity with

Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline or related research field.
  • Minimum 1 year of active research experience in academia, industry, or government labs.
  • Experience with Python, data analysis, simulation, modeling, ML, or scientific computing.
  • Strong understanding of experimental design, hypothesis testing, and statistical interpretation.
  • Familiarity with Git, IDEs, and notebook platforms (Jupyter/Colab).

Responsibilities

  • Design benchmark tasks inspired by real-world scientific workflows.
  • Develop reference solutions using Python and notebooks with rigorous methods.
  • Define evaluation standards to distinguish sound reasoning.
  • Review AI-generated solutions for methodological weaknesses and errors.
  • Collaborate with AI researchers to improve benchmark quality.
  • Refine evaluation methodologies for advanced AI systems.

Skills

Python
Data analysis
Experiment design
Statistical analysis
Scientific computing

Education

Master's degree, PhD, or equivalent practical experience in STEM

Tools

Git
Jupyter
Google Colab
IDE

Job description

Weekday AI is hiring a researcher for a fully remote, full-time role in the United States. You will design complex benchmark tasks, develop reference solutions in Python, and help uncover methodological flaws in state-of-the-art AI models.

The position emphasizes rigorous experimental design, data analysis, and collaboration with AI researchers. Qualifications include a STEM terminal degree or equivalent experience, 1+ year in research, strong Python/data analysis skills, and familiarity with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote Data Science & AI Benchmark Analyst
Remote Data Science & AI Benchmark Analyst

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
AI Benchmark Engineer (Python) — Remote
AI Benchmark Engineer (Python) — Remote

Weekday AI • United States

Remote
USD 83,000 - 124,000
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday AI • United States

Remote
USD 83,000 - 124,000
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote Biology AI Benchmark Scientist (PhD)
Remote Biology AI Benchmark Scientist (PhD)

Weekday AI • United States

Remote
USD 83,000 - 110,000
Remote Physics AI Benchmark Engineer (PhD)
Remote Physics AI Benchmark Engineer (PhD)

Weekday AI • United States

Remote
USD 83,000 - 117,000
Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote Data Science & AI Evaluation Benchmark Designer
Remote Data Science & AI Evaluation Benchmark Designer

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Benchmark Architect (Remote) | Python
AI Benchmark Architect (Remote) | Python

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments