Data Science & Quantitative Analysis Expert

Dorado

United States

Remote

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cincinnatus LLC is hiring for a ground-truth data science role to design, run, and report complex analyses for frontier-model benchmarks. You will create realistic tasks, clean data, and interpret results in clear notebook reports to guide research decisions.

The full-time W-2 position is remote within the United States, approximately 35 hours per week, with collaboration across researchers to ensure rigorous, actionable evaluations.

Qualifications

  • MSc or PhD in statistics, data science, or another quantitative STEM field (or equivalent practical experience).
  • 1+ years in a research, research-engineering, or heavy data-analysis role.
  • Hands-on data-analysis skills: data cleaning, correlations, hypothesis testing, interpretation.

Responsibilities

  • Design realistic data-analysis tasks for evaluation benchmarks.
  • Author notebooks in Jupyter or Colab with reproducible analyses.
  • Compare methods and back findings with spot checks and clear recommendations.
  • Evaluate models’ handling of tasks and validate statistics.
  • Collaborate with researchers to ensure consistent, accurate evaluations.

Skills

Data analysis
Statistical testing
Data cleaning
Python (pandas/NumPy)
Communication of findings

Education

MSc or PhD in statistics/data science/quantitative STEM
Experience in research-heavy analytical domain

Tools

Jupyter Notebooks / Colab
Git

Job description

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced data scientists and quantitative analysts to act as ground-truth experts. You will design complex analysis tasks that simulate real research work — for example, comparing two anomaly-detection algorithms on a dataset, calculating correlations, performing manual spot checks, and summarizing the findings in a notebook clear enough to drive a researcher's decision.

Each task represents one to two days of continuous, focused effort and spans multiple skills: data cleaning, statistical analysis, interpretation, and clear reporting. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fall short on rigorous analytical work.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities
  • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results — based on the kind of analysis you do every day.

  • Author notebooks: Work through your own tasks in Jupyter or Colab, producing clear, reproducible reference analyses.

  • Compare methods: Build tasks that ask for a fair comparison between analytical approaches, backed by spot checks and a clear recommendation.

  • Evaluate models: Review how models handle your tasks, and check whether their statistics and conclusions actually hold up.

  • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate.

3. Core Qualifications
  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.

  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.

  • Deep hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results.

  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.

  • Working proficiency in Python (pandas, NumPy, or similar) and Git.

  • Strong ability to communicate analytical findings in writing for decision-makers.

  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.

  • A perfectionist mindset: high attention to detail, creativity in task design, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Expert - AI/ML
Data Science Expert - AI/ML

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
STEM Researcher — Computational Fields
STEM Researcher — Computational Fields

Dorado • United States

Remote
USD 105,000 - 150,000
QA/Test Engineer
QA/Test Engineer

Dorado • United States

Remote
USD 90,000 - 130,000
Machine Learning Engineer — Model Evaluation & Experimentation
Machine Learning Engineer — Model Evaluation & Experimentation

Dorado • United States

Remote
USD 120,000 - 180,000
Software Domain Expert
Software Domain Expert

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →

Dorado • California (MO), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Red Team Specialist — Failure Modes & Edge Cases
LLM Red Team Specialist — Failure Modes & Edge Cases

Dorado • United States

Remote
USD 120,000 - 180,000
AI Project Coordinator - GenAI
AI Project Coordinator - GenAI

Mercor • San Francisco (CA)

On-site
USD 70,000 - 100,000
AI Project Coordinator - GenAI
AI Project Coordinator - GenAI

Obsidian • San Francisco (CA)

On-site
USD 65,000 - 90,000
AI Project Coordinator - Data Focus
AI Project Coordinator - Data Focus

Obsidian • San Francisco (CA)

On-site
USD 70,000 - 100,000