Senior Data Engineer

EPM Scientific

København

Sur place

DKK 520 000 - 900 000

Plein temps

Il y a 19 heures
Soyez parmi les premiers à postuler
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Résumé du poste

EPM Scientific is seeking an experienced Data Engineer to design, build, and optimize data platforms in Copenhagen. You will work with Azure Databricks, PySpark, Spark SQL, and Delta Lake to deliver scalable data products using a Medallion architecture.

You will implement ingestion frameworks, manage ADLS Gen2 storage, and support CI/CD for data notebooks and pipelines, ensuring data quality, governance, and regulatory compliance across environments.

Qualifications

  • Strong hands-on experience with Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
  • Deep understanding of modern data lakehouse architectures, including Bronze-Silver-Gold (Medallion) design.
  • Experience developing enterprise-scale ETL/ELT pipelines and data integration solutions.
  • Strong knowledge of Azure storage technologies, particularly ADLS Gen2.
  • Experience implementing data governance, lineage, and access control frameworks.
  • Proven experience with data quality, validation, monitoring, and operational support.
  • Experience in regulated environments with compliance and traceability requirements.

Responsabilités

  • Design, develop, and optimize data solutions using Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
  • Develop and maintain data pipelines following the Medallion Architecture to deliver scalable, trusted data products.
  • Implement ingestion frameworks for structured and semi-structured data from APIs and files.
  • Manage and optimize data storage solutions using ADLS Gen2 and related Azure services.
  • Build, schedule, and support workloads using Databricks Workflows/Jobs with Delta Live Tables or Lakeflow as a plus.
  • Develop and maintain data quality controls, validation frameworks, and automated testing.

Connaissances

Azure Databricks
Python/PySpark
Spark SQL
Delta Lake
Medallion Architecture
Data governance
ETL/ELT pipelines
ADLS Gen2
DevOps CI/CD for data

Outils

Delta Lake
Unity Catalog
Delta Live Tables
Lakeflow

Description du poste

  • Design, develop, and optimize data solutions using Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
  • Develop and maintain data pipelines following the Medallion Architecture (Bronze, Silver, Gold) to deliver scalable and trusted data products.
  • Implement ingestion frameworks for structured and semi-structured data from APIs, JSON files, databases, and other enterprise sources.
  • Manage and optimize data storage solutions using Azure Data Lake Storage Gen2 (ADLS Gen2) and associated Azure services.
  • Build, schedule, and support workloads using Databricks Workflows/Jobs, with experience in Delta Live Tables (DLT) or Lakeflow considered a strong advantage.
  • Develop and maintain data quality controls, validation frameworks, reconciliation processes, and automated testing.
  • Design and implement data transformation, cleansing, and standardization processes to support business reporting and analytics.
  • Establish metadata management, data lineage, and governance capabilities using Unity Catalog and related technologies.
  • Develop API-based ingestion and export integrations for upstream and downstream systems.
  • Support solutions involving document metadata management and references to PDF and eLabel assets.
  • Drive performance tuning, partitioning strategies, and platform optimization to improve scalability and efficiency.
  • Implement CI/CD practices for notebooks, code deployments, and automated testing throughout the development lifecycle.
  • Develop monitoring, alerting, error-handling, and reprocessing capabilities to ensure operational stability and reliability.
  • Support regulated data environments with a strong focus on traceability, auditability, reproducibility, and controlled release management.
Key Responsibilities
  • Design, develop, and optimize data solutions using Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
  • Develop and maintain data pipelines following the Medallion Architecture (Bronze, Silver, Gold) to deliver scalable and trusted data products.
  • Implement ingestion frameworks for structured and semi-structured data from APIs, JSON files, databases, and other enterprise sources.
  • Manage and optimize data storage solutions using Azure Data Lake Storage Gen2 (ADLS Gen2) and associated Azure services.
  • Build, schedule, and support workloads using Databricks Workflows/Jobs, with experience in Delta Live Tables (DLT) or Lakeflow considered a strong advantage.
  • Implement schema management practices, including schema enforcement and schema evolution.
  • Develop and maintain data quality controls, validation frameworks, reconciliation processes, and automated testing.
  • Design and implement data transformation, cleansing, and standardization processes to support business reporting and analytics.
  • Establish metadata management, data lineage, and governance capabilities using Unity Catalog and related technologies.
  • Develop API-based ingestion and export integrations for upstream and downstream systems.
  • Support solutions involving document metadata management and references to PDF and eLabel assets.
  • Drive performance tuning, partitioning strategies, and platform optimization to improve scalability and efficiency.
  • Implement CI/CD practices for notebooks, code deployments, and automated testing throughout the development lifecycle.
  • Develop monitoring, alerting, error-handling, and reprocessing capabilities to ensure operational stability and reliability.
  • Support regulated data environments with a strong focus on traceability, auditability, reproducibility, and controlled release management.
Desired Skills and Experience
Required Qualifications
  • Strong hands-on experience with Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
  • Deep understanding of modern data lakehouse architectures, including the Bronze-Silver-Gold (Medallion) design pattern.
  • Experience developing enterprise-scale ETL/ELT pipelines and data integration solutions.
  • Strong knowledge of Azure storage technologies, particularly ADLS Gen2.
  • Experience implementing data governance, lineage, and access control frameworks.
  • Proven experience with data quality, validation, monitoring, and operational support.
  • Experience working in regulated environments where compliance, auditability, and data traceability are key requirements.
Preferred Qualifications
  • Experience with Delta Live Tables (DLT), Lakeflow, and advanced Databricks orchestration capabilities.
  • Familiarity with document-centric data solutions, including PDF and eLabel metadata management.
  • Experience implementing DevOps and CI/CD best practices within data development teams.
  • Experience in pharmaceutical, healthcare, life sciences, or other highly regulated industries.
Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior Staff Software Engineer - Delta
Senior Staff Software Engineer - Delta

Databricks Inc. • Aarhus

Sur place
DKK 900 000 - 1 200 000
Senior Staff Software Engineer - Delta
Senior Staff Software Engineer - Delta

Neon • Aarhus

Sur place
DKK 1 200 000 - 1 800 000
Comprehensive benefits
Diversity and inclusion commitment
Senior Staff Software Engineer - Delta
Senior Staff Software Engineer - Delta

Databricks • Aarhus

Sur place
DKK 800 000 - 1 000 000
Backend Engineer - Event Processing & Intellige...
Backend Engineer - Event Processing & Intellige...

Secomea • København

Hybride
DKK 650 000 - 900 000
Azure Databricks Data Engineer - Medallion Lakehouse
Azure Databricks Data Engineer - Medallion Lakehouse

EPM Scientific • København

Sur place
DKK 520 000 - 900 000
Backend Engineer - Event Processing & Intelligence Services (C# / Python) →
Backend Engineer - Event Processing & Intelligence Services (C# / Python) →

Cph Ai Hub • København

Hybride
DKK 700 000 - 900 000
Backend Engineer - Event Processing & Intelligence Services (C# / Python)
Backend Engineer - Event Processing & Intelligence Services (C# / Python)

Secomea • Region Hovedstaden

Hybride
DKK 600 000 - 900 000
Data engineer
Data engineer

Hamamatsu Photonics A/S • Birkerød

Sur place
DKK 800 000 - 1 000 000
Team Lead Data Engineering - Snowflake / AWS
Team Lead Data Engineering - Snowflake / AWS

EPAM Systems • Danemark

Hybride
DKK 900 000 - 1 200 000
AI & Data Engineer
AI & Data Engineer

Accenture Nordics • København

Sur place
DKK 900 000 - 1 200 000