STEM Researcher - Computational Fields

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.

You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.

Qualifications

  • Master's degree or PhD in a STEM field or equivalent research experience.
  • At least 1 year of experience in an active research role.
  • Experience with Python, data analysis, simulation, ML, or scientific computing.
  • Strong understanding of experimental design and statistical analysis.
  • Familiarity with Git, IDEs, and notebook platforms (Jupyter or Colab).

Responsibilities

  • Design complex benchmark tasks informed by real-world scientific workflows.
  • Develop reference solutions using Python, notebooks, and computational tools.
  • Define evaluation standards to distinguish sound reasoning from flawed conclusions.
  • Review AI-generated solutions and identify methodological weaknesses.
  • Collaborate with researchers to improve benchmark quality and rigor.
  • Refine evaluation methodologies for advanced AI systems.

Skills

Python
Data analysis
Machine learning
Statistical analysis
Scientific computing

Education

Master's degree
PhD

Tools

Git
Jupyter
Google Colab

Job description

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking researchers from computational STEM disciplines—as well as computationally intensive social sciences and humanities—to bring the rigor of real-world research into AI evaluation.

In this role, you will transform scientific methodologies such as experimental design, hypothesis testing, and data-driven analysis into sophisticated, multi-step benchmark tasks that challenge state-of-the-art AI systems. Working closely with AI researchers, you'll help uncover subtle reasoning errors and methodological flaws that only experienced researchers can identify.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities
  • Design complex, research-oriented benchmark tasks inspired by real-world scientific workflows, including study design, experimentation, hypothesis testing, and data analysis.
  • Develop comprehensive reference solutions using Python, notebooks, and computational tools with the rigor expected in professional research.
  • Define clear evaluation standards that distinguish sound scientific reasoning from plausible but incorrect conclusions.
  • Review AI-generated solutions, identifying methodological weaknesses, analytical errors, and flawed reasoning that experienced researchers would recognize immediately.
  • Collaborate with AI researchers and fellow domain experts to improve benchmark quality, consistency, and scientific rigor.
  • Contribute to the continuous refinement of evaluation methodologies for advanced AI systems.
Required Qualifications
  • Master's degree, PhD, or equivalent practical experience in a STEM discipline, computational social science, computational humanities, or another research-intensive field involving programming and data analysis.
  • Minimum 1 year of experience in an active research role within academia, industry, government laboratories, or a similar research environment.
  • Demonstrated experience performing computational research involving Python, data analysis, simulation, modeling, machine learning, or scientific computing.
  • Strong understanding of experimental design, hypothesis testing, statistical analysis, and rigorous interpretation of research findings.
  • Working knowledge of Git, integrated development environments (IDEs), and notebook platforms such as Jupyter or Google Colab.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is preferred.
  • Excellent analytical thinking, attention to detail, creativity, and the ability to solve complex, open-ended problems independently.
  • Strong written communication skills for documenting technical methodologies and research findings.
  • Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
  • Experience designing reproducible computational experiments or research workflows.
  • Familiarity with machine learning, large language models, or AI-assisted research tools.
  • Background in benchmark design, scientific software development, or computational research infrastructure.
  • Experience mentoring researchers, reviewing scientific work, or contributing to peer-reviewed publications.
Why Join
  • Help shape how next-generation AI systems are evaluated using rigorous scientific methodologies.
  • Collaborate with leading AI researchers working on frontier models and advanced evaluation frameworks.
  • Apply your research expertise to improve AI reasoning, reliability, and scientific accuracy.
  • Contribute to impactful work that advances the quality and robustness of AI systems across multiple disciplines.
  • Enjoy the flexibility of a fully remote engagement while working on cutting-edge AI research initiatives.
Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Remote | STEM Expert — $100–$150/hour
Remote | STEM Expert — $100–$150/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 78,000 - 117,000
Clinical Researchers (Remote)
Clinical Researchers (Remote)

Keystone Recruitment • United States

On-site
USD 96,432 - 110,208
Competitive compensation
Flexible scheduling
Collaborate with leading experts
STEM Researchers - Benchmark & First-Author Research
STEM Researchers - Benchmark & First-Author Research

Gramian Consulting Group • United States

Remote
USD 41,000 - 110,000
STEM Researchers - Benchmark & First-Author Research
STEM Researchers - Benchmark & First-Author Research

Gramian Consulting • Massachusetts

Remote
USD 41,328,000 - 110,208,000
Lead author on benchmark paper
Flexible remote work
Biology & Biophysics Research Collaborator (Part-time)
Biology & Biophysics Research Collaborator (Part-time)

Weekday 1 • United States

Remote
USD 110,000 - 152,000
AI Evaluation Specialist
AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
Scientific Computing Research Expert
Scientific Computing Research Expert

Appsierra Group • American Samoa

On-site
USD 96,000 - 110,000