Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics

The Earlham Institute

United Kingdom

Remote

GBP 39,000 - 47,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

The Earlham Institute is inviting applications for a Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics. This 36‑month, full‑time post supports a Moore Foundation project hosted at the Earlham Institute on the Norwich Research Park.

The post holder will lead the computational core of the project: quality control, genome and transcriptome assembly, decontamination and co‑biont separation, and functional/structural annotation across all aims.

Qualifications

  • PhD in bioinformatics, computational biology, genomics, evolutionary biology or closely related discipline
  • Experience analysing large-scale NGS data and de novo genome assembly using long-read sequencing
  • Proficiency in at least one bioinformatics programming language
  • Experience working in Linux/HPC environments
  • Track record of peer‑reviewed publications

Responsibilities

  • Lead genome and transcriptome analysis, including de novo assembly using long‑read data
  • Perform decontamination and co‑biont separation to improve assembly quality
  • Deliver functional and structural annotation of genomes/transcriptomes
  • Perform QC of sequencing data and contribute to data management and FAIR principles
  • Package workflows for public use (e.g., Nextflow/Snakemake, containers) and contribute to reproducible research
  • Disseminate results via publications, conferences and community training

Skills

Bioinformatics
Genomics
Scripting
Linux HPC

Education

PhD in bioinformatics or related

Tools

Galaxy
WorkflowHub
Nextflow
Git

Job description

Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics

Job Title Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics Post Number 1006194 Closing Date 1 Oct 2026 Grade SC6 Starting Salary Salary: £39,000-£46,500

Hours per week 18.5 Project Title Gordon and Betty Moore Foundation – Protist Omics at Scale Expected/Ideal Start Date 02 Nov 2026 Months Duration 36

Job Description
Main Purpose of the Job

The Earlham Institute has been awarded funding by the Gordon and Betty Moore Foundation to develop Protist Omics at Scale, a three-year international methods-development programme. the programme is run in partnership with the Scottish Association for Marine Science, home of the Culture Collection of Algae and Protozoa, and the Bigelow Laboratory for Ocean Sciences. Protists represent the overwhelming majority of eukaryotic diversity, yet they remain hugely under-represented in reference genome databases. Their genomes are large, repetitive, frequently aneuploid and often recovered from mixed or low-biomass cultures, and the resulting data defeat standard assembly and annotation approaches. This project will systematically identify, diagnose and resolve those bottlenecks across three aims: cultured protists with substantial biomass, protists that grow only to low concentrations ,and single-cell genomes and transcriptomes of uncultured protists from the environment. We are looking for a Postdoctoral Research Scientist / Computational Biologist to lead the computational core of the project: quality control, assembly, decontamination and co-biont separation, and structural and functional annotation across all three aims. Where existing tools fail, the postholder will diagnose why and develop what replaces them. A defining requirement of the award is that the resulting methods are usable by the wider research community, so every workflow will be packaged, documented and released openly, with all data deposited under rich, FAIR-compliant metadata. The post offers substantial scope for independent research. This is a rare chance to work at the interface of long-read sequencing, single-cell genomics, and eukaryotic genome assembly, on organisms whose genomes are largely unexplored. The post is based at the Earlham Institute on the Norwich Research Park.

Key Relationships

Internal: Reporting to the Director and Principal Investigator, Prof Neil Hall. The post holder will work closely with the project Co-PIs (Dr Iain Macaulay, Dr Karim Gharbi), the project manager, the postdoctoral scientists and senior research assistants delivering the wet-lab components of the project, and with colleagues across the Institute's Core Bioinformatics, Scientific Computing and Data Infrastructure teams. They will also interact with the Research Faculty Office, the Communications team and the Institute's training group.

External: Co-PIs and researchers at the Scottish Association for Marine Science / CCAP and the Bigelow Laboratory for Ocean Sciences (Single Cell Genomics Center); the Gordon and Betty Moore Foundation and the project's external advisory board; the Darwin Tree of Life Genome Engine and the Wellcome Sanger Institute; the British Society for Protozoology and the wider international protist research community; WorkflowHub, Galaxy, protocols.io, ENA and NCBI; and collaborators and users attending project workshops, hackathons and training events.

Main Activities & Responsibilities

Genome and transcriptome analysis

  • Generate de novo genome assemblies for cultured protists, including chromosome-resolved assemblies using long‑read assemblers integrated with other long‑range information.
  • Develop and apply bespoke assembly strategies tolerant of low‑inputs, amplified material, and of single‑cell genomes, including coloured De Bruijn graph approaches and the adaptation of existing assemblers to single‑cell data.
  • Implement read pooling and co‑assembly strategies across multiple amplified cells of the same population to improve genome recovery.
  • Assemble and evaluate short‑read and long‑read transcriptomes, supporting the analysis of isoform diversity, UTRs, trans‑splicing, non‑canonical splice sites and alternative polyadenylation in protist transcriptomes.

Decontamination, co‑biont separation and chimera removal

  • Remove and/or separate co‑biont and contaminant genomes during assembly clean‑up using the kmer‑ord workflow developed at SAMS and established binning tools (e.g. Eukfinder, CONCOCT, SemiBin2), together with tools under development at EI, adapting them where necessary for single‑cell and low‑input data.
  • Detect and remove physical chimeras introduced during whole‑genome amplification using reference‑based and de novo approaches, and benchmark read‑length thresholds and protocol‑level strategies that minimise chimera formation and data loss.
  • Manually assess and curate assemblies to resolve mis‑assemblies and to separate organellar genomes.

Genome and transcriptome annotation

  • Deliver structural and functional annotation of assembled genomes and transcriptomes, integrating RNA‑seq evidence where available.
  • Assess assembly and annotation quality using BUSCO, k‑mer completeness, contiguity and reference‑based benchmarking, and contribute to the project’s community‑aligned definitions of assembly contiguity and completeness.

Sequencing data quality control and triage

  • Work with lab teams to rapidly analyse data for technical development work. Perform QC of DNA and RNA sequencing libraries and of raw short‑read, long‑read (PacBio and ONT) and long‑range (Hi‑C, Illumina Constellation).
  • Run preliminary assembly of extracted samples to assess DNA/RNA integrity, purity, genome size, heterozygosity and the presence of co‑biont or contaminant material.
  • Feed results back rapidly to the culturing, extraction and library preparation teams at EI and SAMS so that protocols can be iterated, and contribute the evidence underpinning the project’s lineage‑aware extraction and lysis recommendations.

Contribution to the Institute

  • Contribute to the wider scientific life of the Institute, including seminars, group meetings and cross‑platform method development.
  • As agreed with line manager, any other duties commensurate with the nature of the role.

Computational workflow development, packaging and reproducibility

  • Ensure that all computational workflows developed on the project are made publically available.
  • Package analyses reproducibly (e.g. Nextflow/Snakemake, containers, version‑controlled environments) so that methods are portable across the partner laboratories and readily adoptable by the wider community.
  • Support cross‑laboratory benchmarking of workflows between EI, SAMS and the Bigelow SCGC, and troubleshoot deployment on partner infrastructure.

Data management, deposition and protocol dissemination

  • Submit raw data, assemblies, annotations and transcriptome datasets to ENA and/or NCBI, annotated with rich metadata and made findable, accessible, interoperable and reusable (FAIR).
  • Contribute computational protocols and documentation to the project’s protocols.io workspace and project website.
  • Maintain accurate records of analyses and outcomes to support project reporting and milestone tracking.

Dissemination, community engagement and training

  • Prepare results for publication in peer‑reviewed journals and present the work at national and international meetings and & lead in the preparation of papers where appropriate.
  • Contribute to the project’s community engagement programme, including the international launch workshop, the benchmarking hackathon and dedicated training events, and provide hands‑on support to external participants applying project workflows to their own taxa.
  • Contribute technical blog posts and updates to the project website and respond to community feedback on protocols and workflows.
Person Profile
Education & Qualifications

Requirement Importance PhD (awarded, or submitted/close to submission) in bioinformatics, computational biology, genomics, evolutionary biology or a closely related discipline Essential Undergraduate degree in a biological, computational, mathematical or other numerate discipline Essential Formal training or demonstrable self‑directed expertise in eukaryotic genome assembly and annotation Desirable Training in software engineering, reproducible research or research data management Desirable

Specialist Knowledge & Skills

Requirement Importance Practical, hands‑on experience of de novo genome assembly using long‑read data (PacBio HiFi and/or Oxford Nanopore) Essential Proficiency in at least one scripting/programming language used in bioinformatics (e.g. Python, R, Perl, C/C++, Rust) Essential Confident working in a Linux/UNIX command‑line environment and running analyses on an HPC cluster (e.g. Slurm, LSF) Essential Working knowledge of assembly quality assessment, including BUSCO, k‑mer based completeness and contiguity metrics Essential Experience of structural and/or functional annotation of eukaryotic genomes or transcriptomes Essential Competent use of version control (Git/GitHub) and a demonstrable commitment to reproducible, well‑documented analysis Essential Ability to critically evaluate and benchmark competing computational methods and to communicate the results objectively Essential Experience with workflow management systems (e.g. Nextflow, Snakemake, Galaxy) and workflow sharing platforms such as WorkflowHub Desirable Experience of metagenomic or metagenome‑adjacent analysis: read/assembly binning, decontamination and separation of co‑biont genomes Desirable Experience of single‑cell genomics and of the artefacts associated with whole‑genome amplification (chimeras, uneven coverage, allelic dropout) Desirable Experience of scaffolding using long‑range data (Hi‑C, ultra‑long ONT, linked or mapped‑read technologies) Desirable Experience of long‑read transcriptome analysis, isoform discovery and non‑canonical gene structures Desirable Knowledge of protist / microbial eukaryote biology, phylogeny and genome diversity Desirable Familiarity with containerisation and environment management (Docker, Singularity/Apptainer, Conda) Desirable

Requirement Importance Track record of analysing large‑scale next‑generation sequencing datasets from design through to interpretation Essential Publication record in peer‑reviewed journals commensurate with career stage Essential Experience of working with genomic data from non‑model organisms, or with genomes lacking high‑quality references Essential Experience of troubleshooting analyses where the underlying data are imperfect, and of feeding conclusions back to laboratory colleagues to improve upstream protocols Essential Experience of submitting sequence data, assemblies and metadata to public archives (ENA, NCBI/GenBank) Desirable Experience of releasing open‑source software, pipelines or public protocols used by others Desirable Experience of working as part of a multi‑institution or international research consortium Desirable Experience of working alongside a sequencing or genomics service platform Desirable Experience of delivering computational training, workshops or hackathons Desirable

Interpersonal & Communication Skills

Requirement Importance Demonstrated ability to work independently, using initiative and applying problem solving skills Essential Good interpersonal skills, with the ability to work as part of a multidisciplinary team spanning computational and laboratory scientists Essential Excellent communication skills, both written and oral, including the ability to present complex technical information with clarity to non‑specialist audiences Essential Able to plan and prioritise a varied computational workload across three concurrent project aims and to meet agreed deadlines Essential Able to act as an ambassador for the Institute and build key relationships locally on the park and externally Essential Demonstrates a commitment to high personal and professional standards. Assumes responsibility and accountability for the successful completion of projects, assignments or tasks Essential Willingness and ability to support, mentor and train others in computational methods Desirable

Additional Requirements

Requirement Importance Attention to detail Essential Willingness to embrace the expected values and behaviours of all staff at the Institute, ensuring it is a great place to work Essential Commitment to open science, FAIR data principles and the open release of code, workflows and protocols Essential Ability to maintain confidentiality and security of information where appropriate, including in relation to pre‑publication data Essential Ability to undertake occasional national and international travel for project meetings, workshops, hackathons and conferences Essential Willingness to work outside standard working hours when required, including to accommodate meetings across international time zones Essential Promotes equality and values diversity Essential Able to present a positive image of self and the Institute, promoting both the international reputation and public engagement of the Institute Essential

Who We Are

About the Earlham Institute The Earlham Institute harnesses the power of data‑driven biology to accelerate solutions for health, biodiversity, and food security. Based at Norwich Research Park, the Earlham Institute is one of eight institutes strategically funded by BBSRC.

Our science combines world‑class technology, interdisciplinary expertise, and training and development across genomics, engineering biology and data science, to decode the scale and complexity of living systems.

We believe we can achieve more if we work together. That's why we collaborate with the global science community and industry partners, while also inspiring the next generation of scientists and technical specialists.

Our Science Earlham Institute scientists specialise in developing and testing the latest tools and approaches needed to decode living systems and make biological predictions.

We are home to state‑of‑the‑art facilities and technology, creating a unique combination of expertise and infrastructure.

We have dedicated laboratories for genome sequencing, single‑cell analysis, engineering biology, and large‑scale automation; as well as one of the largest supercomputing facilities for life science research in Europe.

Our Advanced Training team also provides access to specialised scientific training to upskill the next generation of research and technical staff.

Our Culture Our collegiate and innovative research environment comes with significant support, including a commitment to your professional development, research and administrative assistance, and opportunities to build collaborations with scientists and industry on the Norwich Research Park, across the UK, and internationally.

The Institute is also home to talented technical and operational staff, whose invaluable contributions enable our science to have the maximum impact. We aim to recognise, reward, and develop all staff and students so that every individual feels able to achieve their best with us.

We work hard to nurture an engaged and positive workplace, centred on core values that include openness, technical excellence, and collaboration. We attract staff from around the world who contribute to - and benefit from - an environment that enables them to deliver world‑class science alongside a supportive and social community.

The Earlham Institute is a BBSRC‑supported research institute on the Norwich Research Park specialising in genomics, single‑cell and spatial biology, and computational biology. The post sits within the Director's group and works across two of the Institute's core science platforms.

The Technical Genomics Group (Dr Karim Gharbi) operates the High‑Throughput Sequencing platform and delivers genome and transcriptome data production, together with technical development for DNA/RNA isolation, library preparation and short‑and long‑read sequencing, including the evaluation of new and emerging methodologies.

The Single‑Cell and Spatial Analysis platform (Dr Iain Macaulay) is one of the most advanced facilities in the UK for single‑cell and spatial genomics of model and non‑model organisms, and led environmental protist sequencing for the Darwin Tree of Life project. It develops imaging and spectral cell sorting approaches and long‑read single‑cell RNA sequencing.

The project team at EI additionally includes a dedicated project manager, experienced protist genomics and single‑cell postdoctoral scientists, and senior research assistants covering DNA/RNA isolation and long‑read library preparation. Externally, the consortium brings together CCAP/SAMS, who maintain one of the world's largest and most taxonomically diverse protist culture collections (~3,200 strains), and the Bigelow Laboratory Single Cell Genomics Center, the world's first single‑cell genomics centre focused on environmental microorganisms.

Postdoctoral Scientist (Bioinformatics) – Protist Genomics The Earlham Institute has been awarded funding by the Gordon and Betty Moore Foundation to develop Protist Omics at Scale, a three‑year international methods‑development programme run in partnership with the Scottish Association for Marine Science, home of the Culture Collection of Algae and Protozoa, and Aalborg University. We are looking for a computational biologist to lead the computational core of the project: quality control, assembly, decontamination and co‑biont separation, and structural and functional annotation across all three aims. Where existing tools fail, the postholder will diagnose why and develop what replaces them. The post is based at the Earlham Institute on the Norwich Research Park. Background: Protists represent the vast majority of eukaryotic diversity but remain significantly under‑represented in reference genome databases. Their genomes are often large, repetitive and genetically complex, and are frequently derived from mixed, low‑biomass or uncultured samples, making them difficult to assemble and annotate using standard genomic approaches. This project aims to address these challenges by systematically identifying and overcoming key bottlenecks in genome and transcriptome assembly from bulk cultures and single cells. Based within the Earlham Institute's Director's Group, the project combines expertise in long‑read sequencing, single‑cell genomics, spatial biology and computational biology. It brings together leading facilities at the Earlham Institute, including the Technical Genomics Group and the Single‑Cell and Spatial Analysis Platform, as well as external collaborators at CCAP/SAMS, home to one of the world's largest protist culture collections, and Aalborg University. The overall objective is to develop and apply innovative methods that enable the generation of high‑quality genomic and transcriptomic resources for previously inaccessible and poorly characterised eukaryotic organisms. The role: This is a postdoctoral computational biology/bioinformatics role focused on developing and applying novel methods for long‑read and single‑cell genome and transcriptome assembly across a diverse range of protist species. The postholder will: • Develop expertise in advanced genome and transcriptome assembly approaches. • Work on complex long‑read and single‑cell sequencing datasets. • Contribute to the development of new computational methods rather than routine analysis. • Develop research software engineering skills, including workflow development, packaging and containerisation. • Create and maintain reproducible bioinformatics workflows using platforms such as Galaxy and WorkflowHub. • Collaborate closely with internal and external partners across the consortium. • Lead or contribute significantly to project outputs. • Publish research findings and present at national and international conferences. • Participate in workshops, hackathons and community training activities. • Support the supervision and development of students where appropriate. The role offers extensive opportunities for career development, networking and collaboration within an internationally recognised genomics research environment. The ideal candidate: The post holder will have, or be close to completing, a PhD in bioinformatics, computational biology, genomics, evolutionary biology or a closely related discipline. They will have practical experience of analysing large‑scale next‑generation sequencing datasets and de novo genome assembly using long‑read sequencing data (PacBio HiFi and/or Oxford Nanopore), together with proficiency in at least one bioinformatics programming language and experience working in a Linux/HPC environment. The successful candidate will have experience of genome or transcriptome analysis, an ability to critically evaluate computational methods, and a track record of contributing to research outputs, including peer‑reviewed publications. Experience of workflow development and reproducible research practices, including version control and workflow management systems, would be advantageous, as would knowledge of single‑cell genomics, protist or microbial eukaryote biology, and software containerisation technologies.

Additional information: This is a full‑time post for a contract of 36 months. Salary on appointment will be within the range £39,000 - £46,500 per annum, depending on qualifications and experience. A starting salary of £40,100 is guaranteed for candidates who can evidence their PhD certificate at appointment; those awaiting confirmation of their PhD award will be appointed at £39,000 until evidence is provided. This role meets the criteria for a visa application, and we encourage all qualified candidates to apply. As a Disability Confident employer, we guarantee to offer an interview to all disabled applicants who meet the essential criteria for this vacancy. The closing date for applications will be 1 October 2026.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics
Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics

Earlham Institute • United Kingdom

Remote
GBP 39,000 - 47,000
Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics
Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics

LLOYD'S REGISTER INTERNATIONAL • Colney

On-site
GBP 39,000 - 47,000
Postdoctoral Scientist (Bioinformatics) - Protist Genomics
Postdoctoral Scientist (Bioinformatics) - Protist Genomics

Earlham Institute • Colney

On-site
GBP 39,000 - 47,000
(Senior) Postdoctoral Research Scientist
(Senior) Postdoctoral Research Scientist

The Earlham Institute • Norwich

Hybrid
GBP 39,000 - 58,000
Postdoctoral Fellow
Postdoctoral Fellow

European Molecular Biology Laboratory • Cambridgeshire and Peterborough

Hybrid
GBP 42,000 - 54,000
Hybrid working patterns
Private medical insurance
30 days annual leave
+2
Postdoctoral Fellow
Postdoctoral Fellow

European Bioinformatics Institute • Cambridgeshire and Peterborough

Hybrid
GBP 42,000 - 48,000
Hybrid working patterns
Visa and relocation support
Career development resources
Galaxy Workflow Expert
Galaxy Workflow Expert

The Earlham Institute • Norwich

Hybrid
GBP 38,000 - 47,000
Postdoctoral Fellow
Postdoctoral Fellow

European Molecular Biology Laboratory (EMBL) • Hinxton

Hybrid
GBP 49,000 - 66,000
Hybrid working patterns
Private medical insurance
30 days annual leave
+1
DevOps Engineer (BioFAIR Methods Commons)
DevOps Engineer (BioFAIR Methods Commons)

LLOYD'S REGISTER INTERNATIONAL • Norwich

On-site
GBP 47,000 - 53,000
Bioinformatician - Pathogen
Bioinformatician - Pathogen

Ellison Institute of Technology Oxford • Oxford

Hybrid
GBP 50,000 - 75,000
Travel allowance
Bonus
Enhanced holiday pay
+8