Biomedical Subject Matter Expert

Mind Moves

United States

Remote

USD 179,000 - 220,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mind Moves is seeking a Biomedical Subject Matter Expert to provide senior scientific guidance on evaluation framework development, benchmark dataset validation, and scientific review for the NLM Biomedical AI Challenges.

The role centers on mapping dbGaP variables to standard vocabularies (LOINC, UMLS, CUIs), curating semantic annotations, and ensuring interoperability via the dbGaP FHIR API and related data dictionaries. Remote within the US preferred.

Qualifications

  • Advanced degree (MS or PhD preferred) in bioinformatics, biomedical informatics, epidemiology, genetics, library/information science, or a related field — or equivalent professional experience.
  • Demonstrated experience with biomedical controlled vocabularies and terminology standards (e.g., UMLS, LOINC, MeSH, PhenX, or OBO Foundry ontologies such as HPO or OBI).
  • Familiarity with clinical/genomic data repositories and data-sharing models comparable to dbGaP (e.g., NIH CDE Repository, TOPMed, other GWAS/genomic consortia).
  • Experience with semantic web or ontology tooling (e.g., OWL, RDF, Protégé).
  • Experience with Python or R.
  • Strong written communication skills for documenting mapping rationale.
  • Working knowledge of information-retrieval and classification evaluation metrics used to grade automated systems (precision, recall, F1-score, NDCG) and how to translate expert judgment into a defensible ground-truth/reference dataset.
  • Preferred Qualifications: Prior experience with dbGaP data dictionaries, submission packets, or the dbGaP FHIR API.
  • Experience with data-dictionary harmonization tools (e.g., D2Refine or similar).
  • Familiarity with GWAS/genomic consortium data harmonization processes.
  • Experience with FHIR resource modeling in a research-data or genomics context.
  • Prior experience building or contributing to benchmark/gold-standard datasets for an evaluation campaign or shared task (e.g., TREC-style relevance judgments, NLP/IR bake-offs) and calibrating multiple annotators to a shared rubric.

Responsibilities

  • Map dbGaP variables to controlled vocabularies and common data elements, including PhenX measurement protocols, using the dbGaP data dictionary's VARIABLE_SOURCE and SOURCE_VARIABLE_ID fields, and classify mappings by confidence level (e.g., identical, comparable, or related) consistent with dbGaP/PhenX conventions.
  • Curate and validate semantic annotations linking dbGaP variables to standard terminologies such as the UMLS Metathesaurus, LOINC, and UMLS Concept Unique Identifiers (CUIs), following the approach used by groups like NLM's Lister Hill Center, Medical Data Models, and NHLBI's TOPMed program.
  • Support FHIR-based interoperability by ensuring annotated vocabularies are correctly represented in the dbGaP FHIR schema and accessible via the dbGaP FHIR API.
  • Design and curate the expert-adjudicated reference ("gold standard") mapping set against which Track 1 participant submissions are scored.
  • Author and maintain annotation guidelines and adjudication protocols so reference-set decisions are reproducible and defensible under review or challenge.
  • Run or oversee inter-annotator agreement checks across SME reviewers and resolve disagreements before a mapping enters the reference set.
  • Advise on Track 2 relevance judgments — labeling dbGaP studies as relevant/not relevant to natural-language research queries.
  • Contribute to held-out or adversarial "probe" cases used to detect memorized or hardcoded submissions during anti-gaming review.

Skills

Biomedical vocabularies
Genomic data repositories
Python or R
Written communication
Evaluation metrics
dbGaP experience
GWAS data harmonization
FHIR modeling
Benchmark datasets
Semantic web/ontology tooling

Education

MS/PhD in bioinformatics or related

Tools

OWL/RDF/Protégé
D2Refine

Job description

About Mind Moves

Mind Moves is a women-owned Washington, D.C.-based firm that helps government and business partners navigate digital transformation using "human-in-the-loop" AI. Our expert team has delivered responsibly developed AI products that drive millions in impact across agencies like the National Institutes of Health (NIH).

Program Overview

The National Library of Medicine (NLM) is launching Biomedical AI Challenges — a structured prize competition program designed to catalyze innovation in AI-powered tools for biomedical research infrastructure. The program seeks to improve AI-enhanced semantic search across PubMed/PMC and establish foundational data interoperability by mapping raw, cohort-specific dbGaP study variables directly to precise, standardized international vocabularies/ontologies (such as LOINC, RxNorm, SNOMED CT, and the UMLS Metathesaurus).

Why This Work Matters

Right now, dbGaP study variables and metadata are described in disparate, cohort-specific shorthand rather than standardized codes. Because standardized terminologies attach stable, unambiguous codes to clinical concepts, they are what make it possible for different studies, systems, and AI tools to "talk" about the same phenotype, lab result, or exposure in the same way. Without that shared vocabulary layer, cross-study comparison stays manual and error-prone, dbGaP's rich phenotypic data remains difficult to discover, and researchers can spend months pursuing a controlled-access request only to find the cohort doesn't match their needs. The SME's terminology mapping and curation work is the foundation this entire Challenge is built on: it is what turns free-text variable descriptions into the standardized, computable concepts that both Challenge tracks — and, ultimately, the broader research community — depend on for reliable semantic search and cross-study interoperability.

Position Summary

The Biomedical Subject Matter Expert (SME) serves as a senior scientific advisor and technical authority to provide expert guidance on evaluation framework development, benchmark dataset validation, and scientific review. The Biomedical SME contributes meaningfully to shaping the scientific integrity and rigor of challenge-related tasks.

Key Responsibilities
  • Map dbGaP variables to controlled vocabularies and common data elements, including PhenX measurement protocols, using the dbGaP data dictionary's VARIABLE_SOURCE and SOURCE_VARIABLE_ID fields, and classify mappings by confidence level (e.g., identical, comparable, or related) consistent with dbGaP/PhenX conventions.

  • Curate and validate semantic annotations linking dbGaP variables to standard terminologies such as the UMLS Metathesaurus, LOINC, and UMLS Concept Unique Identifiers (CUIs), following the approach used by groups like NLM's Lister Hill Center, Medical Data Models, and NHLBI's TOPMed program.

  • Support FHIR-based interoperability by ensuring annotated vocabularies are correctly represented in the dbGaP FHIR schema and accessible via the dbGaP FHIR API.

Reference Set & Evaluation Design
  • Design and curate the expert-adjudicated reference ("gold standard") mapping set against which Track 1 participant submissions are scored, including explicit criteria for the ONT_NONE discard class (variables with no valid target-vocabulary concept).

  • Author and maintain annotation guidelines and adjudication protocols so reference-set decisions are reproducible and defensible under review or challenge.

  • Run or oversee inter-annotator agreement checks across SME reviewers and resolve disagreements before a mapping enters the reference set.

  • Advise on Track 2 relevance judgments — labeling dbGaP studies as relevant/not relevant to natural-language research queries — in a form suitable for computing ranking metrics.

  • Contribute to held-out or adversarial "probe" cases used to detect memorized or hardcoded submissions during anti-gaming review.

Required Qualifications
  • Advanced degree (MS or PhD preferred) in bioinformatics, biomedical informatics, epidemiology, genetics, library/information science, or a related field — or equivalent professional experience.

  • Demonstrated experience with biomedical controlled vocabularies and terminology standards (e.g., UMLS, LOINC, MeSH, PhenX, or OBO Foundry ontologies such as HPO or OBI).

  • Familiarity with clinical/genomic data repositories and data-sharing models comparable to dbGaP (e.g., NIH CDE Repository, TOPMed, other GWAS/genomic consortia).

  • Experience with semantic web or ontology tooling (e.g., OWL, RDF, Protégé).

  • Experience with Python or R.

  • Strong written communication skills for documenting mapping rationale.

  • Working knowledge of information-retrieval and classification evaluation metrics used to grade automated systems (precision, recall, F1-score, NDCG) and how to translate expert judgment into a defensible ground-truth/reference dataset.

Preferred Qualifications
  • Prior experience working directly with dbGaP data dictionaries, submission packets, or the dbGaP FHIR API.

  • Experience with data-dictionary harmonization tools (e.g., D2Refine or similar).

  • Familiarity with GWAS/genomic consortium data harmonization processes.

  • Experience with FHIR resource modeling in a research-data or genomics context.

  • Prior experience building or contributing to benchmark/gold-standard datasets for an evaluation campaign or shared task (e.g., TREC-style relevance judgments, NLP/IR bake-offs) and calibrating multiple annotators to a shared rubric.

Position Type: Part-Time Contract (Independent Contractor)

Program: NLM Biomedical AI Challenges Program

Period of Performance: Sept 2026 – March 2027

Project Hours: 500 – 600 total hours

Compensation: $130-160/hr

Location: Remote within the US

Interview Process

We expect the process to include a short hiring exercise followed by 1 20-30 min interview for final candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Biomedical Ontology & Data Standards Expert
Remote Biomedical Ontology & Data Standards Expert

Mind Moves • United States

Remote
USD 179,000 - 220,000
Biomedical Data Annotator | Remote | US
Biomedical Data Annotator | Remote | US

Mind Moves • Northern (KY)

Hybrid
USD 34,000 - 48,000
Remote Biomedical Informatics SME - 58854
Remote Biomedical Informatics SME - 58854

Turing • Boston (MA)

On-site
USD 83,000 - 124,000
Fully remote working environment
Work on cutting-edge AI projects
Potential for contract extension
Remote Biomedical Informatics SME - 58854
Remote Biomedical Informatics SME - 58854

Turing • Los Angeles (CA)

On-site
USD 100,000 - 134,000
Work in a fully remote environment
Opportunity to work on cutting-edge AI projects
Potential for contract extension based on performance
Remote Biomedical Informatics SME - 58854
Remote Biomedical Informatics SME - 58854

Turing • Denver (CO)

On-site
USD 83,000 - 165,000
Fully remote work environment
Work on cutting-edge AI projects
Potential for contract extension based on performance
+1
Remote Biomedical Informatics SME - 58854
Remote Biomedical Informatics SME - 58854

Turing • San Francisco (CA)

On-site
USD 110,000 - 193,000
Work in a fully remote environment
Opportunity on cutting-edge AI projects
Potential for contract extension
+3
Bioinformatics Engineer | Remote | US
Bioinformatics Engineer | Remote | US

Mind Moves • Northern (KY)

Hybrid
USD 138,000 - 179,000
Remote Biomedical Informatics SME - 58854
Remote Biomedical Informatics SME - 58854

Turing • Austin (TX)

On-site
USD 83,000 - 152,000
Fully remote work environment
Work on cutting-edge AI projects
Potential for contract extension
+1
Freelance Subject Matter Expert – Internal Medicine
Freelance Subject Matter Expert – Internal Medicine

Cactus Communications • Houston (TX)

On-site
USD 180,000 - 260,000
Project-based work (no fixed hours)
Senior Bioinformatics Scientist - III
Senior Bioinformatics Scientist - III

AllSTEM Connections • Cambridge (MA)

On-site
USD 198,000 - 242,000
Benefits