Job Title: Senior Bioinformatics Scientist III
Location: Cambridge, MA 02141
100% onsite
Duration: 2 Years Contract
Shift: Monday Thru Friday (8 AM to 5 PM) - 40 Hours/Week
Position Summary:
The Precision Genetics group within the Data and Genome Sciences Department is seeking a skilled Contractor to join our Computational Precision Genetics team. We are looking for a data scientist who combines deep expertise in human genetic data analysis with strong machine learning capabilities and hands-on experience integrating multi-omics data, to support our target identification, patient stratification, and biomarker discovery efforts.
Key Responsibilities:
- Data Ingestion: Query and harmonize external resources to acquire relevant genetic, genomic, and multi-omics datasets (e.g., dbSNP, 1000 Genomes Project, gnomAD, GTEx, Ensembl, Open Targets, ClinVar, GWAS Catalog, UK Biobank, Gene Expression Omnibus).
- Genetic/Genomic Data Analysis: Perform quality control (QC) and analysis of genetic/genomic data, including genotype imputation from array data, variant calling and annotation using state-of-the-art methods (e.g., IMPUTE, Minimac, Eagle, BEAGLE, GATK, bcftools, samtools, ANNOVAR, VEP).
- Statistical Genetics: Conduct genetic association analyses at scale, including GWAS/PheWAS, rare-variant burden and collapsing tests, fine-mapping, colocalization, polygenic scores, and Mendelian randomization (e.g., PLINK, REGENIE, SAIGE, GCTA, SuSiE, coloc, LDSC).
- QTL Analysis: Conduct QTL analysis to identify genetic loci associated with quantitative and molecular traits, including eQTL, sQTL, and pQTL mapping, utilizing tools such as tensorQTL, FastQTL, PLINK, or R/qtl.
- Population Genetics Analysis: Analyze genetic variation across populations, including allele frequency estimation, linkage disequilibrium, relatedness, and ancestry/population structure analysis.
- Machine Learning: Develop, benchmark, and validate machine learning models on high-dimensional genetic and molecular data for tasks such as variant effect prediction, patient stratification, and biomarker or treatment-response prediction; apply rigorous cross-validation, control for batch and ancestry confounding, and use interpretability methods to translate models into testable biological hypotheses (e.g., scikit-learn, XGBoost, PyTorch, SHAP).
- Multi-Omics Data Integration: Integrate genetic datasets with other omics layers, including transcriptomic (bulk and single-cell RNA-seq), epigenomic, proteomic (e.g., OLINK, mass spectrometry), and spatial data, to provide comprehensive insights into gene function and disease biology (e.g., DESeq2, limma, Seurat, scanpy).
- Documentation and Reproducibility: Prepare detailed documentation of analysis methods and results in a timely manner, and deliver version-controlled, reproducible analysis workflows (e.g., Git, Nextflow, Snakemake).
Required Qualifications:
- Ph.D. in Genetics, Genomics, Statistical Genetics, Computational Biology, or a related field.
- A proven track record of over 5 years in genetic data analysis.
- Strong understanding of statistical methods and genetic data analysis and integration (e.g., variant analysis, GWAS and QTL mapping, population genetics, genomic annotations).
- Demonstrated experience applying machine learning to high-dimensional biological data, including feature engineering, model selection, validation, and avoidance of overfitting and confounding.
- Hands-on experience integrating multi-omics data (e.g., transcriptomics, proteomics, epigenomics) with genetic data.
- Proficiency in R, Python, and Bash, with the ability to establish best practices for reproducible data analyses.
- Experience with high-performance computing (HPC) systems and AWS Cloud Computing (e.g., IAM, S3 buckets).
- A collaborative and self-motivated individual with a strong work ethic, capable of managing multiple objectives in a dynamic environment and adapting to changing priorities.
- Excellent written and verbal communication skills.
Preferred Qualifications:
- Experience with real-world and large-scale biobank genetic data (e.g., UK Biobank, All of Us, FinnGen, electronic health record-linked cohorts).
- Experience with deep learning approaches for genomics, including sequence-based and variant-effect prediction models.
- Familiarity with single-cell and spatial transcriptomics analysis.
- Experience supporting drug target identification and validation, or biomarker discovery in a pharmaceutical or biotechnology setting.
- Familiarity with workflow managers (e.g., Nextflow, Snakemake) and containerization (e.g., Docker, Singularity).
- Proficient in human genetic/human genomic data analysis tools and techniques-e.g., variant analysis, population genetics, genomic annotations.
- Multi-Omics Data Integration
- Proficiency in R, Python, and Bash.
- High-performance computing (HPC) systems and AWS Cloud Computing (e.g., IAM, S3 buckets).