Biostatistician - AI Trainer

DataAnnotation

Indiana (PA)

Remote

USD 55,000 - 172,000

Full time

18 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible, remote and project-based
Competitive compensation

Job summary

DataAnnotation is seeking experienced computational biologists and bioinformaticians to design realistic AI-driven analysis tasks, run them through frontier agents, and grade outputs against professional standards.

You'll write clear grading rubrics and explain why a result passes or fails, highlighting issues like batch effects, inappropriate tests, or misinterpreted data. Remote, project-based work with multinational collaboration.

Qualifications

  • Must have 3+ years hands-on experience analyzing real biological data in industry, academia, or research lab.
  • Proficient in writing, debugging, and explaining analysis code in Python and/or R (at least one; both preferred).
  • Deep knowledge in at least one core area (genomics, transcriptomics, proteomics/metabolomics, genetics, ML in biology).
  • Experience end-to-end with multistep analyses from raw data to interpretation.

Responsibilities

  • Design realistic analysis tasks from your workflows, including prompts and necessary files.
  • Run tasks through frontier AI agents and grade outputs against professional standards.
  • Write clear grading rubrics and justify pass/fail decisions with evidence.
  • Flag concrete failures with evidence, such as ignored batch effects or inappropriate statistical tests.

Skills

Genomics knowledge
Machine learning in biology
Data interpretation

Education

Master's or PhD in biology/biostatistics/computer science

Tools

Python
R

Job description

We’re looking for experienced computational biologists and bioinformaticians to help train AI models. You’ll design realistic analysis tasks from your own practice, run them through frontier AI systems, and grade what comes back against a professional standard.

The models can talk fluently about biological data analysis. What they can’t yet do reliably is the real work: QC an RNA-seq count matrix, prioritize variants, diagnose a batch effect, or audit a pipeline. Your judgment of when a result is real becomes the benchmark those models are measured against.

What you’ll actually do
  • Design realistic analysis tasks from your own workflows: the scenario, the prompt, and the files an agent would need (count matrices, sample sheets and metadata, VCFs, pipeline logs and QC output, analysis notebooks).
  • Run tasks through frontier AI agents and grade the deliverable (an analysis report, annotated table, figure set, or notebook) against the standard you’d hold a colleague to.
  • Write clear grading rubrics (the right normalization, the right multiple-testing correction, the right biological reading) and explain why a response passes or fails.
  • Flag concrete failures with evidence: batch effects and confounders ignored, wrong statistical test or uncorrected p-values, misused reference version, silently dropped samples, or code that doesn’t do what the narrative claims.

Problems draw on whatever you know best:

  • Genomics and variant interpretation.
  • Bulk and single-cell transcriptomics.
  • Proteomics, metabolomics, and multiomics integration.
  • Population and statistical genetics.
  • Machine learning applied to biology.
Roles this fits

Common backgrounds: Bioinformatics Scientist, Genomics Data Scientist, Computational Biologist.

What we look for
  • 3+ years hands-on analyzing real biological data in industry, an academic lab, or a research institute (counted after undergraduate study).
  • You write, debug, and can explain your own analysis code in Python and/or R (at least one required, both preferred).
  • Depth in at least one of: genomics and variant interpretation; bulk or single-cell transcriptomics; proteomics / metabolomics / multiomics; population and statistical genetics; ML applied to biology; clinical genomics; metagenomics or phylogenetics.
  • You’ve owned a multistep analysis end to end, from raw or messy data to final interpretation, and can judge whether a result is real (normalization, batch effects, multiple-testing correction, biological interpretation).
  • Master’s or PhD (or current PhD candidate) in biology, bioinformatics, biostatistics, computer science, or a related field, completed in the U.S., Canada, Europe, or the UK.
  • Clear written English, comfort with ambiguity, and familiarity with AI/LLM tools like Claude or ChatGPT.
  • Based in the United States, Canada, or the UK (Ireland and Australia may also be accepted).
Compensation

Up to $40 – $125+/hr depending on task difficulty and specialization. Many contributors add $10k–$100k+ a year; some make it their full-time income.

About DataAnnotation

DataAnnotation is where 100k+ experts train the world’s leading AI models. $150M+ paid to contributors to date, and the average contributor stays 5+ years. Flexible, remote, and always project-available.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Biostatistician - AI Trainer
Biostatistician - AI Trainer

DataAnnotation • California (MO)

Remote
USD 55,000 - 172,000
Epidemiologist - AI Trainer
Epidemiologist - AI Trainer

DataAnnotation • Illinois

Remote
USD 55,000 - 172,000
Epidemiologist - AI Trainer
Epidemiologist - AI Trainer

DataAnnotation • Missouri

Remote
USD 55,000 - 172,000
Pharmacologist - AI Trainer
Pharmacologist - AI Trainer

DataAnnotation • Minnesota

Remote
USD 55,000 - 172,000
Biomedical Researcher - AI Trainer
Biomedical Researcher - AI Trainer

DataAnnotation • Hawaii

Remote
USD 55,000 - 172,000
Remote work
Biomedical Researcher - AI Trainer
Biomedical Researcher - AI Trainer

DataAnnotation • Alaska

Remote
USD 55,000 - 172,000
Flexible schedule
Remote work
Project-based work
Neuroscientist - AI Trainer
Neuroscientist - AI Trainer

DataAnnotation • Iowa (LA)

Remote
USD 55,000 - 172,000
Flexible schedule
Remote work
Pharmacologist - AI Trainer
Pharmacologist - AI Trainer

DataAnnotation • Idaho

Remote
USD 55,000 - 172,000
Flexible schedule
Remote work
Project-based assignments
Pharmacologist - AI Trainer
Pharmacologist - AI Trainer

DataAnnotation • United States

Remote
USD 55,000 - 172,000
Neuroscientist - AI Trainer
Neuroscientist - AI Trainer

DataAnnotation • United States

Remote
USD 55,000 - 172,000