Genetics Data Engineer 1

Biopharma Careers

Bengaluru

On-site

INR 2,400,000 - 4,200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Biopharma Careers is seeking a Senior Principal Data Engineer - Genetics to drive a scalable data engineering layer for quantitative genetics. You will design and implement reproducible pipelines, containerized workflows, and data governance practices to ensure high-quality data for downstream analyses.

You will collaborate with geneticists, data scientists, and IT teams, deploying tools into TREs like All of Us and UK Biobank equivalents, while optimizing performance and enabling rapid genetic

Qualifications

  • Bachelor's or Master’s degree in computer science, data engineering, bioinformatics, or a related field.
  • Minimum 6 years relevant experience in data engineering, with exposure to healthcare, Medical or biological data.
  • Track record of building and deploying pipelines and solutions that show your stakeholders what is possible beyond what they had imagined or requested.
  • Strong expertise in Python and ideally R for data ingestion, processing, cleaning, and pipeline orchestration.
  • Experience with large, complex datasets in cloud or HPC environments, including tools such as Spark, S3, and cloud-native compute platforms (AWS, GCP, Azure).
  • Experience with containerization (Docker, Singularity) and infrastructure-as-code practices for reproducible deployments in secure environments.
  • Solid understanding of data modeling, versioning, and reproducibility principles.
  • Experience with methods and requirements for medical or genetic data privacy, including data governance and controlled-access data handling.

Responsibilities

  • Advance genetics data engineering to enable downstream analysis.
  • Collaborate with scientists to bring data/tools into TREs and deploy.
  • Ingest and maintain connections to genetic reference databases and integrate with the internal knowledge graph.
  • Automate data QC for diverse genomic data types.
  • Develop reproducible analysis pipelines using workflow managers and containerized environments.
  • Return results to internal platforms per each biobank's privacy and data protection policies.
  • Optimize query performance and pipeline execution for rapid assessments and in-licensing due diligence.
  • Contribute to the design and implementation of agentic AI workflows for automated genetic evidence generation.
  • Build and maintain dashboards and data services that expose genetic evidence to project teams.
  • Link genetic data to AI tools and platforms, ensuring seamless data flow between genetic analyses and downstream decision-support systems.

Skills

Python
R
Data engineering
Cloud computing
Containerization
Workflow tools
APIs
Data governance
Genomics data
Communication

Education

Bachelor's/Master's in CS or bioinformatics

Tools

Spark
AWS
GCP
Azure
S3
Nextflow
Snakemake
Hail
PLINK
bcftools
samtools
Docker
Singularity
Plotly Dash
Shiny

Job description

Work Your Magic with us!

Ready to explore, break barriers, and discover more? We know you’ve got big plans – so do we! Our colleagues across the globe love innovating with science and technology to enrich people’s lives with our solutions in Healthcare, Life Science, and Electronics. Together, we dream big and are passionate about caring for our rich mix of people, customers, patients, and planet. That's why we are always looking for curious minds that see themselves imagining the unimaginable with us.

Senior Principle Data Engineer - Genetics
Your Role

You will advance our human quantitative genetics strategy by providing the data engineering fundamentals that enable downstream analysis. You will collaborate with quantitative geneticists, data scientists, data engineers, platform experts, IT, and others to ensure that data availability and quality are never the bottlenecks for our analyses. In that collaboration, you will provide the vision and implementation for how our FAIR data environment works across internal and external platforms such as biobank trusted research environments (TREs).

You will work with other scientists to:
  • Bring software tools, reference datasets, and genetic data into internal environments.
  • Bring tools, containers, and reference data into TREs (e.g., UK Biobank Research Analysis Platform, All of Us Researcher Workbench) and manage their deployment and versioning.
  • Ingest and maintain connections to genetic reference databases (OpenTargets, GWAS Catalog, ClinVar, dbSNP, OMIM, HGMD, ChEMBL, DrugBank) and integrate them with the internal knowledge graph (Synaptix) and analytics platforms.
  • Automate data QC for diverse genomic data types.
  • Develop, test, and execute reproducible analysis pipelines using workflow managers and containerized environments (Docker, Singularity).
  • Return results from TREs to our internal platforms in accordance with each biobank's privacy and data protection policies.
  • Optimize query performance and pipeline execution to support rapid-turnaround target assessments and in-licensing due diligence (20-25 targets per year requiring fast genetic evaluation).
  • Contribute to the design and implementation of agentic AI workflows for automated genetic evidence generation, integrating genetics pipelines with the broader agentic AI platform.
  • Build and maintain interactive dashboards and data services that expose genetic evidence to project teams, leadership, and due diligence committees.
  • Link genetic data to our AI tools and platforms, ensuring seamless data flow between genetic analyses and downstream decision-support systems. Experience building production-grade data pipelines for genetic and genomics datasets,including familiarity with common formats ( VCF, PLINK/BED/BIM/FAM/BGEN,GWAS summary statistics).
  • Automate routine analyses including standard safety assessments, target-disease association lookups, and genetic evidence reports to minimize geneticist time on repetitive tasks.
  • Manage cloud compute budgets and optimize resource usage within TREs to maximize analytical throughput within allocated funding.
Who You Are

You have substantial expertise in data engineering for scientific and genomic data and are comfortable working both on strategic questions as well as hands-on implementation. You have

  • Bachelor's or Master’s degree in computer science, data engineering, bioinformatics, or a related field.
  • Minimum 6 years relevant experience in data engineering, With exposure to healthcare, Medical or biological data.
  • Track Record of building and deploying pipelines and solutions that show your stakeholders what is possible beyond what they had imangined or requested.
  • Strong expertise in Python and Ideally R for data ingestion, processing, cleaning, and pipeline orchestration.
  • Expertise in working with large, complex datasets in cloud or HPC environments, including tools such as Spark, S3, and cloud-native compute platforms (AWS, GCP, Azure). Experince with both tabular data and free text processing.
  • Experience with containerization (Docker, Singularity) and infrastructure-as-code practices for reproducible deployments in secure environments.
  • Solid understanding of data modeling, versioning, and reproducibility principles.
  • Experience with methods and requirements for medical or genetic data privacy, including data governance and controlled-access data handling.
Preferred Qualifications
  • Experience with additional biobank platforms or multi-ethnic datasets (e.g., Biobank Japan, Galatea, FinnGen).
  • Hands on experience with biobank trsuted research enviornments such as UK Biobank ( DNAnexus), All of Us Researcher workbench, or similar platforms.
  • Experience with genomis-specific tools and frameworks ( Hail, PLINK, bcftools, samtools, liftover) and workflow managers ( Nextflow, WDL, Snakemake).
  • Experience building dashboards or visualization tools (e.g., Shiny, Streamlit, Plotly Dash) for scientific audiences.
  • Familiarity with agentic AI Frameworks or experience building LLM-integrated data pipelines and automated reporting tools.
  • Experience with API development (REST/GraphQL) and MCP for serving analytical results to downstream applications.
What we offer:

We are curious minds that come from a broad range of backgrounds, perspectives, and life experiences. We believe that this variety drives excellence and innovation, strengthening our ability to lead in science and technology. We are committed to creating access and opportunities for all to develop and grow at your own pace. Join us in building a culture of inclusion and belonging that impacts millions and empowers everyone to work their magic and champion human progress!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Data Engineer
Sr Data Engineer

Illumina • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Bioinformatics Scientist
Bioinformatics Scientist

Biopharma Careers • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Data Engineer 2
Data Engineer 2

Illumina • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Computational Biology Scientist
Computational Biology Scientist

ThinkBio.Ai, Inc. • Ernakulam

On-site
INR 1,800,000 - 2,400,000
Senior Computational Biologist
Senior Computational Biologist

ThinkBio.Ai, Inc. • Ernakulam

On-site
INR 1,800,000 - 3,200,000
Specialist , Data Engineering
Specialist , Data Engineering

MSD Malaysia • Hyderabad

Hybrid
INR 1,800,000 - 3,000,000
Senior Manager, Data Engineering
Senior Manager, Data Engineering

Gen • Pune District

On-site
INR 1,800,000 - 3,000,000
Flexible working options
Competitive pay
Well-being programs
Data Engineer, Translational Data Management, Automation & AI
Data Engineer, Translational Data Management, Automation & AI

Biopharma Careers • Hyderabad

On-site
INR 3,200,000 - 5,200,000
Senior Manager, Data Engineering
Senior Manager, Data Engineering

Gen • Chennai District

On-site
INR 1,500,000 - 2,500,000
Flexible working options
Competitive pay
Well-being programs
Computational Biology & Bioinformatics Scientist
Computational Biology & Bioinformatics Scientist

Zifo • Chennai District

On-site
INR 800,000 - 1,200,000