Senior Data Pipeline Engineer/Developer

mydnainc

Houston (TX)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Gene by Gene is seeking a Senior Data Pipeline Engineer/Developer to design, build, and own scalable data pipelines moving genomic and operational data from lab instruments to analytics and clinical reporting.

In this Python-first, contractor role, you will set technical direction for the data platform, ensure data quality, lineage, and observability, and guide existing teams through implementation while adhering to regulatory and security requirements.

Qualifications

  • 8+ years in software or data engineering.
  • 4+ years building production data pipelines.
  • 3+ years production cloud experience (AWS preferred).
  • Experience in regulated environments; HIPAA/GxP audits.
  • BS in CS or related field; advanced degree preferred.

Responsibilities

  • Pipeline engineering and integration across LIMS and lab data.
  • Architect data pipelines and establish standards.
  • Own data quality, lineage, and observability.
  • Ensure regulatory compliance and security governance.

Skills

SQL proficiency
Python proficiency
Workflow orchestration
Docker & Kubernetes
Data quality & tests
Communication
Genomics data formats
AWS data stack
Spark / Kafka / Kinesis
dbt & Snowflake
Infrastructure as code
LIMS integration

Education

Bachelor's degree in CS or related
Advanced degree preferred

Tools

Argo Workflows
Dagster
Airflow
Terraform
Pulumi
Nextflow

Job description

Our Purpose

Our mission is to build a healthier and more connected world with precision health and genealogy services. We empower individuals with actionable insights into their genetic makeup, fostering a deeper understanding of their ancestry, health, and wellness. By integrating the experience of Gene by Gene Laboratory Services, FamilyTreeDNA genealogy, and myDNA reporting services, we strive to deliver cutting-edge genetic testing and personalized solutions that inspire informed decisions and enhance quality of life. Our team is dedicated to advancing the field of genomics through innovation, research, and a commitment to excellence.

Our Values

All employees are expected to demonstrate our values of Innovate, One Team, and Integrity when carrying out the accountabilities and responsibilities of their role. This is how we show up every day for ourselves, our colleagues and our customers and strategic partners to deliver our vision and strategic goals.

Position Overview

We are seeking a Senior Data Pipeline Engineer/Developer to design, build, and own the data pipelines that move genomic and operational data from lab instruments and LIMS through to analytics, products, and clinical reporting. In this senior individual-contributor role you will set technical direction for our data platform, build production pipelines that are reliable and reproducible at scale, and provide architectural and code-level guidance to existing engineering teams. You will own data quality, lineage, and observability end to end, and operate in a regulated environment where reproducibility and auditability are non-negotiable. This is a Python-first, full-stack engineering role. This is a contractor position, slated for a 12-month term.

Accountabilities and Responsibilities
  • Pipeline Engineering & Integration
    • Builds and operates production ETL and ELT for high-volume genomic data as well as operational and business data.
    • Integrates data across LIMS, lab instruments, internal applications, and third-party sources through robust, well-tested interfaces.
    • Develops across the stack in Python (data services, internal APIs, and supporting application code) and provides architectural and code-level guidance to engineering teams on data-layer integration.
  • Architecture & Technical Leadership
    • Sets technical direction for batch and streaming data pipelines by evaluating frameworks, orchestration, storage, and processing patterns, making recommendations, and leading adoption.
    • Models and tunes data stores (Microsoft SQL Server and PostgreSQL, plus a cloud warehouse or lake) for performance and scale.
    • Defines and enforces engineering standards for testing, CI/CD, infrastructure as code, code review, and architecture decision records.
  • Data Quality & Pipeline Observability
    • Owns data quality, lineage, and observability, including freshness, completeness, schema-drift detection, cost-per-job, and SLAs.
    • Builds pipelines for reproducibility and audit-readiness, incorporating versioned data and code, lineage, decision logging, access controls, and evidence collection.
    • Partners with security and compliance on data privacy, PII and PHI handling, and regulatory requirements across the data lifecycle.
  • Regulatory Compliance and Security Governance
    • Apply deep understanding of CAP/CLIA, HIPAA, GDPR, and GxP regulations specifically to data pipeline architecture.
    • Oversee protected health information (PHI) handling, data lineage, retention, and comprehensive audit logging.
    • Produce and maintain the critical operational evidence required to carry the data platform successfully through compliance audits.
    • Adhere to strict data-governance controls necessary for securely handling sensitive genomic data across global regions.
    • Enforce least-privilege access and ensure zero data egress to personal or unapproved infrastructure.
    • Utilize exclusively de-identified or synthetic data within development environments.
    • Maintain strict compliance with international data-residency and localization requirements.
Position Requirements
  • Skills and Knowledge
    • Strong SQL proficiency on Microsoft SQL Server and PostgreSQL, including schema design, query tuning, and performance troubleshooting.
    • Strong Python proficiency across the stack (data pipelines, backend services, and APIs). This is a Python-first role.
    • Production experience with a workflow orchestrator (Argo Workflows, Step Functions, Prefect, Dagster, Airflow, or similar).
    • Production experience with containers (Docker) and Kubernetes, which the lead orchestrator (Argo Workflows) runs on.
    • Strong testing discipline, including unit, integration, and data-quality or contract tests.
    • Excellent written and verbal communication, including the ability to explain data tradeoffs to technical and compliance stakeholders.
    • Genomics or NGS data formats and handling (FASTQ, BAM/CRAM, VCF).
    • Bioinformatics workflow engines (Nextflow, WDL/Cromwell, or Snakemake).
    • AWS data stack proficiency (S3, Glue, EMR, Batch, Lambda, Redshift) and/or AWS HealthOmics.
    • Distributed processing and streaming frameworks (Spark, Kafka, Kinesis).
    • Data warehouse and transformation tooling (Snowflake, dbt).
    • Infrastructure as code practices (Terraform, CDK, CloudFormation, or Pulumi).
    • Familiarity with .NET (C#) and/or C++ as supporting languages for integrating with existing services.
    • Knowledge of LIMS integration and on-premises plus hybrid data architectures.
  • Experience
    • 8+ years of professional software or data engineering experience.
    • 4+ years building and operating production data pipelines at scale.
    • 3+ years of production cloud experience (AWS preferred).
    • Demonstrable experience working in regulated environments. Compliance is a hard requirement for this role.
    • Experience supporting HIPAA or GxP audits.
  • Education
    • Bachelor's degree in Computer Science, Bioinformatics, or a related field, or equivalent professional experience.
    • An advanced degree in a quantitative field is preferred.
Why Join Us

At Gene by Gene, you’ll join a mission-driven team advancing the science of genetics and discovery. You’ll have the opportunity to shape meaningful campaigns, tell compelling brand stories, and collaborate with talented professionals who share your passion for creativity, curiosity, and impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Pipeline Engineer/Developer
Senior Data Pipeline Engineer/Developer

Gene by Gene • Houston (TX)

On-site
USD 110,000 - 170,000
Senior Data Pipeline Engineer/Developer
Senior Data Pipeline Engineer/Developer

myDNA, Inc. • Houston (TX)

On-site
USD 140,000 - 190,000
Sr. Clinical Bioinformatics Scientist
Sr. Clinical Bioinformatics Scientist

myDNA • Houston (TX)

On-site
USD 90,000 - 120,000
Senior Genomics Data Pipeline Engineer
Senior Genomics Data Pipeline Engineer

mydnainc • Houston (TX)

On-site
USD 140,000 - 190,000
Senior Genomics Data Pipeline Engineer
Senior Genomics Data Pipeline Engineer

myDNA, Inc. • Houston (TX)

On-site
USD 140,000 - 190,000
Senior Genomics Data Pipeline Architect
Senior Genomics Data Pipeline Architect

Gene by Gene • Houston (TX)

On-site
USD 140,000 - 180,000
Senior Genomics Data Pipeline Architect
Senior Genomics Data Pipeline Architect

Gene by Gene • Houston (TX)

On-site
USD 110,000 - 170,000
Implementation Specialist
Implementation Specialist

mydnainc • Houston (TX)

Hybrid
USD 110,000 - 160,000
Medical & Health Savings
Dental & Vision coverage
401(k) Matching
+2
Lead Bioinformatics Engineer
Lead Bioinformatics Engineer

Quest Diagnostics • Baltimore (MD)

Hybrid
USD 130,000 - 160,000
Health benefits package
401(k) with company match
Annual incentive plans
+1
Senior Software Engineer, .NET
Senior Software Engineer, .NET

GeneDx • United States

On-site
USD 165,000 - 175,000
Paid Time Off (PTO)
Health, Dental, Vision and Life insurance
401k Retirement Savings Plan
+2