Data Science Expert - AI/ML

Obsidian

San Francisco (CA)

On-site

USD 100,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cincinnatus LLC is recruiting experienced data scientists and quantitative analysts to design ground-truth tasks for frontier AI models. This full-time W-2 role is fully remote in the United States, approx. 35 hours per week, with tasks ranging from data cleaning to reporting in notebooks.

You will work closely with researchers to ensure rigorous analyses and clear recommendations for decision-makers. The role requires 1+ year in research-heavy analytics, strong Python (pandas/NumPy), Jupyter or

Qualifications

  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.
  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.
  • Hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and interpretation.
  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working proficiency in Python (pandas, NumPy) and Git.
  • Strong ability to communicate analytical findings in writing for decision-makers.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • Perfectionist mindset with high attention to detail and independent work in open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results.
  • Author notebooks: Work through tasks in Jupyter or Colab, producing clear, reproducible analyses.
  • Compare methods: Build tasks for fair comparisons between analytical approaches with spot checks.
  • Evaluate models: Review how models handle tasks, ensuring statistics hold up.
  • Work as a team: Compare notes with researchers to keep evaluations consistent and accurate.

Skills

Python
Pandas/NumPy
Data cleaning
Statistical analysis
Git
Notebook reporting

Education

MSc/PhD in statistics or data science
1+ years research/data-analysis experience

Tools

Jupyter
Colab

Job description

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced data scientists and quantitative analysts to act as ground-truth experts. You will design complex analysis tasks that simulate real research work — for example, comparing two anomaly-detection algorithms on a dataset, calculating correlations, performing manual spot checks, and summarizing the findings in a notebook clear enough to drive a researcher's decision.

Each task represents one to two days of continuous, focused effort and spans multiple skills: data cleaning, statistical analysis, interpretation, and clear reporting. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fall short on rigorous analytical work.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities
  • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results — based on the kind of analysis you do every day.

  • Author notebooks: Work through your own tasks in Jupyter or Colab, producing clear, reproducible reference analyses.

  • Compare methods: Build tasks that ask for a fair comparison between analytical approaches, backed by spot checks and a clear recommendation.

  • Evaluate models: Review how models handle your tasks, and check whether their statistics and conclusions actually hold up.

  • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate.

3. Core Qualifications
  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.

  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.

  • Deep hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results.

  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.

  • Working proficiency in Python (pandas, NumPy, or similar) and Git.

  • Strong ability to communicate analytical findings in writing for decision-makers.

  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.

  • A perfectionist mindset: high attention to detail, creativity in task design, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Dorado • United States

Remote
USD 90,000 - 120,000
Machine Learning Engineer — Model Evaluation & Experimentation
Machine Learning Engineer — Model Evaluation & Experimentation

Dorado • United States

Remote
USD 120,000 - 180,000
STEM Researcher — Computational Fields
STEM Researcher — Computational Fields

Dorado • United States

Remote
USD 105,000 - 150,000
QA/Test Engineer
QA/Test Engineer

Dorado • United States

Remote
USD 90,000 - 130,000
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →

Dorado • California (MO), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Domain Expert - Fully Remote
Software Domain Expert - Fully Remote

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Domain Expert
Software Domain Expert

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
LLM Red Team Specialist — Failure Modes & Edge Cases
LLM Red Team Specialist — Failure Modes & Edge Cases

Dorado • United States

Remote
USD 120,000 - 180,000
AI Project Coordinator - Data Focus
AI Project Coordinator - Data Focus

Obsidian • San Francisco (CA)

On-site
USD 70,000 - 100,000
AI Project Coordinator - Data Focus
AI Project Coordinator - Data Focus

Mercor • San Francisco (CA)

On-site
USD 80,000 - 110,000