Data Engineer - Data Platform & Ontology (m/f/d)

Statista

Hamburg, Berlin

Hybrid

EUR 70.000 - 100.000

Vollzeit

Vor 8 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Statista in Hamburg is seeking a Data Engineer on our Healthcare Platform to own data ingestion, entity resolution, and platform infrastructure end-to-end. You will design and operate scalable pipelines, harden orchestration, and ensure system reliability, security, and cost efficiency.

A central challenge is building our unified healthcare data ontology, covering hospitals, departments, metrics, and classifications.

Qualifikationen

  • 3+ years in data engineering building production pipelines and data platforms.
  • Experience with healthcare domain knowledge is a plus.
  • Familiarity with hospital structures and medical classifications is valuable.
  • Ability to design scalable, reproducible data workflows.

Aufgaben

  • Pipeline infrastructure and orchestration to ingest data from diverse sources and lake storage.
  • Drive entity resolution and master data management for international hospital entities.
  • Establish data contracts and schema governance for reproducible datasets.
  • Optimize storage, queries, and compute costs in AWS and Snowflake.
  • Automate testing and deployment workflows for data pipelines using GitHub Actions.
  • Collaborate with analytics engineers and data scientists to deliver production-ready datasets.

Kenntnisse

Data Ingestion
Pipeline Orchestration
Python
Analytical SQL
REST APIs
Cloud Lakes
Entity Resolution
MDM
Data Governance
CI/CD Automation
GitHub Actions

Ausbildung

Bachelor's or Master's in Computer Science, Data Science, Software Engineering

Tools

Prefect
Airflow
Dagster
dbt contracts
Pydantic
JSON Schema
Snowflake
AWS S3
Terraform
OpenMetadata
DataHub

Jobbeschreibung

At Statista, we're all about facts and data, for we are the world's leading business data platform. By providing reliable and easy-to-use data as well as various data analytics products and services, we empower people worldwide to make fact-based decisions.

Founded in Hamburg in 2007, we have quickly grown into a global company with offices in major cities such as London, New York, Berlin and Tokyo. And we still have a lot of plans. Our constant growth does not only prove our success, but also keeps creating new development and career opportunities for our employees.

We value and celebrate our diverse culture. You are welcome here for who you are, no matter where you come from, what you look like, or whether you prefer bar graphs to pie charts. Your story matters - keep writing it as part of our team.

Are you ready to join us?

About the Role

As Data Engineer on our Healthcare Platform, you will own the foundational data ingestion, entity-resolution, and platform infrastructure end-to-end. You will design and operate scalable batch/streaming data pipelines, harden platform orchestration, and ensure system reliability, security, and cost efficiency.

A central challenge of this role is building and operationalizing our unified healthcare data ontology—a consistent semantic model covering hospitals, departments, specialties, metrics, and classification standards. Working closely with Analytics Engineers, Data Scientists, and Methodology experts, you will turn heterogeneous, multi-country hospital data into a coherent, highly queryable data asset.

Key Responsibilities
  • Pipeline Infrastructure & Orchestration: Build, optimize, and operate reliable ELT pipelines (using Python, SQL, and Prefect/Airflow) to ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage (S3, Apache Iceberg).
  • Healthcare Ontology & Entity Resolution: Drive the implementation of entity resolution and master data management (MDM) for international hospital entities, mapping raw source data to canonical structures and maintaining standardized vocabularies (e.g., ICD/OPS, specialty taxonomies).
  • Data Contracts & Schema Governance: Establish strict data contracts (Pydantic, dbt contracts) and schema management, ensuring dataset reproducibility, data lineage tracking, and automated validation across all platform pipelines.
  • Platform Efficiency & Cloud Infrastructure: Optimize data storage, query execution, and compute costs across AWS and Snowflake, keeping data assets performant, secure, and cost-effective.
  • Automation & CI/CD: Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Infrastructure as Code (Terraform).
  • Cross-Functional Data Enablement: Partner directly with Analytics Engineers, Data Scientists, and domain experts to deliver documented, research-grade, and production-ready datasets.
Qualifications
Core Requirements (Must-Haves)
  • Data Ingestion & Pipeline Orchestration: Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes (S3/Iceberg). Hands-on experience with modern orchestrators (Prefect, Airflow, or Dagster).
  • Entity Resolution & Data Governance: Practical experience with entity resolution/record linkage frameworks (e.g., Splink, dedupe, recordlinkage) and schema management/data contracts (Pydantic, dbt contracts, or JSON Schema).
  • Cloud Platform & Warehouse Infrastructure: Deep hands-on experience in an AWS production environment (S3, ECS/EC2) combined with cloud data warehouses (Snowflake).
  • Automated CI/CD & Workflow Automation: Proven track record of automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions.
Nice-to-Haves (What Will Make You Stand Out)
  • Knowledge Graphs & Healthcare Terminologies: Exposure to ontology/semantic frameworks (RDF/OWL, SKOS, Neo4j, LinkML) or international medical classifications/vocabularies (SNOMED CT, ICD/OPS, FHIR).
  • Metadata & Lineage Tooling: Experience operating metadata registries and lineage catalogs (e.g., OpenMetadata, DataHub, dbt docs).
  • Infrastructure as Code (IaC): Proficiency in using Terraform to declaratively manage cloud resources and environments.
Your Profile
  • Degree: Bachelor's or Master's in Computer Science, Data Science, Software Engineering, or a related quantitative field.
  • Experience: 3+ years in data engineering building production pipelines and data platforms; including a sustained period within one organization seeing a core platform or product through build launch iteration.
  • Domain Knowledge: Healthcare domain experience is a plus (basic understanding of healthcare KPIs, quality metrics, or benchmarking concepts; familiarity with hospital structures and medical classification systems like ICD/OPS is especially valuable).
  • Mindset: Strong analytical and systems mindset, with a proven ability to transform ...
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Consultant - Healthcare Data (m/f/d)
Consultant - Healthcare Data (m/f/d)

Remotely • Hamburg

Vor Ort
EUR 70.000 - 100.000
Hybrid work
International team
Mobility subsidy
+5
Consultant Data Science - Healthcare (m/f/d)
Consultant Data Science - Healthcare (m/f/d)

Statista • Berlin, Hamburg

Hybrid
Work from abroad up to 30 days a year
Hybrid work and flex-time
International team and social events
+3
Consultant Data Science - Healthcare (m/f/d)
Consultant Data Science - Healthcare (m/f/d)

Remotely • Hamburg

Vor Ort
EUR 70.000 - 100.000
Work from abroad up to 30 calendardays
Hybrid work and flex-time
International team and social events
+5
Consultant - Healthcare Data (m/f/d)
Consultant - Healthcare Data (m/f/d)

Statista • Berlin, Hamburg

Hybrid
EUR 70.000 - 110.000
Work from abroad up to 30 calendar 
-
Hybrid work and flex-time
International team and social events
+5
Consultant Data Science - Healthcare (m/f/d)
Consultant Data Science - Healthcare (m/f/d)

Statista Inc. • Deutschland

Hybrid
EUR 70.000 - 95.000
Work from abroad
Hybrid work
International team
+5
Senior Full-Stack Engineer - Data Intelligence Applications (m/f/d)
Senior Full-Stack Engineer - Data Intelligence Applications (m/f/d)

Statista • Hamburg, Berlin

Hybrid
EUR 120.000 - 170.000
Hybrid work model
International team
Mobility and fitness options
+2
Consultant - Healthcare Data (m/f/d)
Consultant - Healthcare Data (m/f/d)

Meyandy LLC • Hamburg, Berlin

Vor Ort
USD 69.000 - 104.000
Data & Analytics Engineer (m/f/d)
Data & Analytics Engineer (m/f/d)

Meyandy LLC • Hamburg, Berlin

Vor Ort
USD 81.000 - 115.000
Hybrid work
International team
Career & training
+2
Data & Analytics Engineer (m/f/d)
Data & Analytics Engineer (m/f/d)

Statista • Berlin, Hamburg

Hybrid
EUR 70.000 - 110.000
Hybrid work and flex-time
International team and events
Subsidized urban mobility
+3
Data & Analytics Engineer (m/f/d)
Data & Analytics Engineer (m/f/d)

Statista • Deutschland

Hybrid
EUR 70.000 - 100.000
Hybrid work & flex-time
Career & training opportunities
Mental health support