Data Engineer (Bioinformatics)

Remote Worker LTD.

United States

Remote

USD 82,000 - 119,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive base salary
Pension scheme
Holiday allowance
Home & remote working
Wellbeing support

Job summary

Our Future Health is seeking a Data Engineer with a strong background in bioinformatics and genetics to build data pipelines supporting the Clinical Research Recruitment Service. You will combine genomics, bioinformatics and modern data engineering to work with large-scale data and contribute to research that improves prevention, detection and treatment.

You will collaborate with scientists, software engineers and product managers to design reusable, scalable production pipelines, applying best

Qualifications

  • Experience building robust, scalable data pipelines for large datasets.
  • Strong Python development and data wrangling skills.
  • Knowledge of genomic data formats (VCF, BGEN) and QC tools.
  • Cloud experience (Azure preferred) and distributed computing.
  • Familiarity with GA4GH and FAIR data principles.
  • Experience in an agile software delivery team.

Responsibilities

  • Support reusable data pipelines to identify clinical trial participants.
  • Translate data requirements into robust ETL transformations.
  • Develop prototypes for complex pipelines using existing workflows.
  • Adopt best practices: code reviews, unit tests, clean code.
  • Collaborate with scientists, software engineers and product managers.
  • Build scalable production pipelines for genomic data.

Skills

Python
Data pipelines
Bioinformatics
Agile
Stakeholder management
Collaboration

Tools

Docker
Kubernetes
Nextflow
WDL/Cromwell
Airflow
Prefect
Dagster
Git/GitHub
Spark
Databricks
Parquet
VCF
BGEN
PLINK
bcftools
QCtools

Job description

We’re looking for a Data Engineer with a strong background in bioinformatics and genetics to help build the data pipelines supporting Our Future Health’s growing Clinical Research Recruitment Service.

This is an opportunity to combine genetics, bioinformatics and modern data engineering, working with data at significant scale while contributing directly to research that could improve how diseases are prevented, detected and treated.

Our Future Healthis an ambitious collaboration between the public, charity and private sectors, designed tohelp people live healthier lives for longer through better prevention, earlierdetectionand improved treatment of diseases. We will speed up the discovery of new methods of early disease detection, and the evaluation of new diagnostic tools, to helpidentifyand treat diseases early,when outcomes are usually better. With over 2.7M volunteers across the UK, we’re now the world’s biggest health research programme of its kind, and our volunteer group is also more diverse than other, similar health research programmes.

Technology and data are central to our mission. Our systems power web sites, clinics across the UK, secure analytics and research systems, pipelines that process highly sensitive health and genetic data, and we are continuing to grow our engineering capability to support this ambition.

Our Clinical Research Recruitment Service helps life sciences and academic organisations identify potential participants for clinical studies. As the service grows, we need to turn scientific workflows and manual processes into robust, reusable and scalable production pipelines.

We’re looking for a Data Engineer with a solid understanding and experience of bioinformatics, in particular tools and methods associated with genomic data. You can design, build and test pipelines using a range of different technologies. You know how to create repeatable and reusable products and can communicate to and between technical and non-technical stakeholders, with the ability to facilitate discussions and manage different perspectives within a multidisciplinary team including scientists, software engineers, product managers and other data engineers.

Essential Duties and Responsibilities
  • Support the build of re-usable data pipelines used to identify prospective clinical trial participants.
  • Produce logic for data transformation steps as code, which meets the requirements for our end users and builds well curated, accessible and quality controlled data for analysis.
  • Developing prototypes for pipelines for complex transformations drawing on existing workflows developed in industry and academia.
  • Keep abreast of best practice in data engineering across industry, research and Government and facilitating the adoption of standards.
  • Providing technical input into the upstream parts of the data pipeline, including the specification and transfer of data from data providers.
  • Routine ad-hoc data curation activities requiring hands on development of bespoke ETL cleaning scripts using languages such as Python.
  • Working with researchers to understand the data requirements and work with them to deliver the data needed for their projects.
Requirements

We welcome applications from all who may not feel they match the full criteria, so if you have most of the below, we'd like to hear from you:

  • Experience building and maintaining robust, scalable and efficient data pipelines. Capable of processing very large amounts of data based on feeds from multiple systems using a range of different technologies.
  • Can listen to the needs of technical and business stakeholders and interpret them, and effectively manage stakeholder expectations.
  • Detailed knowledge and understanding of genomic data (experience in genotyping and imputation is advantageous).
  • Experience using bioinformatics file standards (VCF, BGEN etc) and tools (PLINK, bcftools, QCtools etc)
  • Highly proficient in Python.
  • Highly proficient in version control and Git/GitHub.
  • Experience of workflow management tools, e.g. Nextflow, WDL/Cromwell, Airflow, Prefect, Dagster
  • Understanding of containerisation (e.g. Docker) and deployment (e.g. Kubernetes).
  • Good understanding of cloud environments (ideally Azure), distributed computing and scaling workflows and pipelines
  • Understanding of common data transformation and storage formats, e.g. Apache Parquet.
  • Awareness of data standards such as GA4GH ( https://www.ga4gh.org/ ) and FAIR ( https://www.go-fair.org/fair-principles/ ).
  • Experience with Spark, Databricks, data lakes.
  • Follow best practices like code reviews, clean code and unit tests.
  • Experience working in an agile development team.
Why join us?

You’ll be joining at an important point as we scale our clinical research recruitment capabilities and move from manual scientific workflows towards reusable engineering solutions.

That means the opportunity to build things that don’t exist today, tackle challenging problems with huge datasets and see a clear connection between your engineering work and real-world health research.

We’re looking for someone pragmatic, collaborative and comfortable working through ambiguity, someone who enjoys solving difficult problems rather than simply maintaining established systems.

It’s a rare combination of science, engineering, scale and purpose.

Hiring process

We feel hiring should be transparent and give you a real sense for what Our Future Health is like. Here's what you can expect:

  • Initial chat with our Talent team (30 min) to get to know each other, discuss the role, and answer any questions you have.
  • 1st interview with our Head of Data Engineering (30 min) - this is an opportunity to learn more about each other and align on role expectations.
  • Technical interview with 2-3 members from our engineering team (60 min). This will include a short task you’ll need to prepare in advance of the interview and is designed to get a sense for how you'd approach the real-world responsibilities of this role.
  • Final stage competency interview (60 min) with a small cross‑functional panel of your potential new colleagues. This will focus on how you collaborate, work with stakeholders, and navigate ambiguity and challenges in real‑world projects.
Benefits
  • Competitive base salary from £62,000 per annum
  • Generous Pension Scheme – We invest in your future with employer contributions of up to 12%
  • 30 Days Holiday pro rata + Bank Holidays – Enjoy a generous holiday allowance with the flexibility to take bank holidays when it suits you
  • Enhanced Parental Leave – Supporting you during life’s biggest moments
  • Cycle to Work Scheme – Save 25-39% on a new bike and accessories through salary sacrifice
  • Home & Tech Savings – Get up to 8% off on IKEA and Currys products, spreading the cost over 12 months through salary sacrifice
  • EV Scheme – Save up to 40% on a brand new electric vehicle all-inclusive package through salary sacrifice
  • £1,000 Employee Referral Bonus – Know something amazing? Get rewarded for bringing them on board!
  • Wellbeing Support – Access to Mental Health First Aiders, plus 24/7 online GP services and an Employee Assistance Programme for you and your family
  • A Great Place to Work – We have a lovely Central London office in Holborn, and offer flexible and remote working arrangements
You may work remotely from anywhere within the UK, with occasional travel to London.

At Our Future Health, we recognise the importance of having a diverse workforce and ensuring that all candidates, regardless of their background, have equitable access to our application process. We proactively encourage applicants who identify as having a disability, neurodiversity, or long-term health conditions to let us know if they require any reasonable adjustments as part of their application process.

If you do require any reasonable adjustments, please email us at talent@ourfuturehealth.org.uk

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Remote Worker LTD. • United States

Remote
USD 98,000 - 119,000
Competitive base salary
Generous pension scheme
30 days holiday + bank holidays
+7
Staff Data Engineer
Staff Data Engineer

TAU Ventures • Massachusetts

On-site
USD 200,000 - 325,000
Staff Data Engineer
Staff Data Engineer

Iterative Health • New York (NY), Cambridge (MA)

On-site
USD 200,000 - 325,000
Lead Scientist
Lead Scientist

Genomics • North Carolina

On-site
USD 120,000 - 180,000
Competitive Salary
Career Path
Continuous Learning
+2
Lead Scientist
Lead Scientist

Genomics • Raleigh (NC), Durham (NC)

On-site
USD 120,000 - 170,000
25 days leave
Pension scheme
Private health insurance
Data Engineer, Product
Data Engineer, Product

Futurhealth • United States

Hybrid
USD 120,000 - 180,000
Fully remote work policy
Unlimited Paid Time Off
13 company holidays + floating holiday
+7
Informatics Engineer
Informatics Engineer

Prime Medicine • Cambridge (MA)

On-site
USD 134,000 - 163,000
Equity
Health, dental, vision
401(k) match
+1
Data Engineer - Sigma Team - Austin
Data Engineer - Sigma Team - Austin

Biorce • Austin (TX)

On-site
USD 110,000 - 145,000
Hybrid work model
Private health coverage
MacBook provided
+1
Clinical Data Engineer
Clinical Data Engineer

Biorce • Austin (TX)

On-site
USD 110,000 - 160,000
Hybrid work model
MacBook provided
Comprehensive health benefits
+2
Scientist
Scientist

Genomics plc • Triangle (VA)

On-site
USD 100,000 - 180,000
Life insurance
Pension
Group income protection
+4