Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.
Statista — глобальная платформа деловых данных. Разрабатываете, оптимизируете и эксплуатируете ELT‑пайплайны на Python, SQL, Prefect/Airflow для разноязычных источников. Работаете с S3, Iceberg, Snowflake и AWS, внедряете схемы и контракты данных.
Требуется опыт конфигурации и автоматизации тестирования пайплайнов через GitHub Actions, знание международных словарей и онтологий, и свободное владение английским языком. Гибридная работа, командировки и развитие в глобальном масштабе.
Описание: Statista is a global business data platform that provides reliable, easy-to-use data, analytics products, and services to support fact-based decision-making worldwide. Its Healthcare Platform transforms international hospital data into structured, queryable healthcare data assets.
Build, optimize, and operate reliable ELT pipelines using Python, SQL, and Prefect/Airflow; Ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage including S3 and Apache Iceberg; Drive entity resolution and master data management for international hospital entities; Map raw source data to canonical structures and maintain standardized vocabularies such as ICD/OPS and specialty taxonomies; Establish data contracts and schema management with Pydantic and dbt contracts; Ensure dataset reproducibility, data lineage tracking, and automated validation across platform pipelines; Optimize data storage, query execution, and compute costs across AWS and Snowflake; Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Terraform; Partner with Analytics Engineers, Data Scientists, and domain experts to deliver documented, research-grade, production-ready datasets.
Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes; Hands-on experience with Prefect, Airflow, or Dagster; Practical experience with entity resolution or record linkage frameworks; Experience with schema management and data contracts using Pydantic, dbt contracts, or JSON Schema; Deep hands-on experience in an AWS production environment, including S3 and ECS/EC2; Experience with cloud data warehouses, including Snowflake; Proven experience automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions; 3+ Years of experience in data engineering building production pipelines and data platforms; Experience taking a core platform or product through build, launch, and iteration within one organization; Bachelor's or Master's degree in Computer Science, Data Science, Software Engineering, or a related quantitative field; Strong analytical and systems mindset; Ability to transform messy, heterogeneous international data into clean, well-governed, highly structured data assets; Fluent English; Highly structured, curious, detail-oriented, and collaborative working style; Nice to have: Knowledge of ontology or semantic frameworks, international medical classifications or vocabularies, metadata registries and lineage catalogs, Terraform, healthcare domain experience, German.
Work from abroad up to 30 calendar days a year; Hybrid work and flex-time; International team and social events; Subsidized urban mobility and access to fitness and wellness options; Free access to Langdock; Career and training opportunities; Attractive locations and modern offices; Mental health support with OpenUp; Some benefits apply only to the German entity and to Junior-level roles or above.