We’re looking for experienced computational biologists and bioinformaticians to help train AI models. You’ll design realistic analysis tasks from your own practice, run them through frontier AI systems, and grade what comes back against a professional standard.
The models can talk fluently about biological data analysis. What they can’t yet do reliably is the real work: QC an RNA-seq count matrix, prioritize variants, diagnose a batch effect, or audit a pipeline. Your judgment of when a result is real becomes the benchmark those models are measured against.
What you’ll actually do
- Design realistic analysis tasks from your own workflows: the scenario, the prompt, and the files an agent would need (count matrices, sample sheets and metadata, VCFs, pipeline logs and QC output, analysis notebooks).
- Run tasks through frontier AI agents and grade the deliverable (an analysis report, annotated table, figure set, or notebook) against the standard you’d hold a colleague to.
- Write clear grading rubrics (the right normalization, the right multiple-testing correction, the right biological reading) and explain why a response passes or fails.
- Flag concrete failures with evidence: batch effects and confounders ignored, wrong statistical test or uncorrected p-values, misused reference version, silently dropped samples, or code that doesn’t do what the narrative claims.
Problems draw on whatever you know best:
- Genomics and variant interpretation.
- Bulk and single-cell transcriptomics.
- Proteomics, metabolomics, and multiomics integration.
- Population and statistical genetics.
- Machine learning applied to biology.
Roles this fits
Common backgrounds: Bioinformatics Scientist, Genomics Data Scientist, Computational Biologist.
What we look for
- 3+ years hands-on analyzing real biological data in industry, an academic lab, or a research institute (counted after undergraduate study).
- You write, debug, and can explain your own analysis code in Python and/or R (at least one required, both preferred).
- Depth in at least one of: genomics and variant interpretation; bulk or single-cell transcriptomics; proteomics / metabolomics / multiomics; population and statistical genetics; ML applied to biology; clinical genomics; metagenomics or phylogenetics.
- You’ve owned a multistep analysis end to end, from raw or messy data to final interpretation, and can judge whether a result is real (normalization, batch effects, multiple-testing correction, biological interpretation).
- Master’s or PhD (or current PhD candidate) in biology, bioinformatics, biostatistics, computer science, or a related field, completed in the U.S., Canada, Europe, or the UK.
- Clear written English, comfort with ambiguity, and familiarity with AI/LLM tools like Claude or ChatGPT.
- Based in the United States, Canada, or the UK (Ireland and Australia may also be accepted).
Compensation
Up to $40 – $125+/hr depending on task difficulty and specialization. Many contributors add $10k–$100k+ a year; some make it their full-time income.
About DataAnnotation
DataAnnotation is where 100k+ experts train the world’s leading AI models. $150M+ paid to contributors to date, and the average contributor stays 5+ years. Flexible, remote, and always project-available.