Mid Data Engineer

Jobtailor

Barcelona

Presencial

EUR 70.000 - 100.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Jobtailor is seeking a data engineer to own and operate production data pipelines on Databricks, with strong PySpark/SQL expertise. You will migrate legacy Hive Metastore to Unity Catalog and maintain Iceberg sharing with Snowflake, while building dbt models and Airflow DAGs.

You will work closely with client stakeholders and the team lead to optimize performance and cost, ensuring data quality and reliable incident handling in a fast-paced environment.

Formación

  • 3+ years operating production data pipelines.
  • Proficient in PySpark and SQL to read, debug, and modify existing pipelines.
  • Experience with Delta Lake MERGE/upsert, table properties, and partitioning.
  • Knowledge of Databricks Workflows, cluster config, and job troubleshooting.
  • Familiarity with Unity Catalog catalogs, schemas, grants, lineage, and metastore model.
  • Experience with Snowflake warehouses, roles and grants, and general operating model.
  • Ability to read and optimize Snowflake queries and manage compute costs.
  • Experience with dbt models, sources, tests, and incremental materializations.
  • Understanding of dbt project structure and deployment workflow.
  • Experience building and maintaining Airflow DAGs and orchestration.
  • Handling retries, backfills, and idempotent task design.
  • Strong SQL including window functions, joins, and reading transformation logic.
  • Python scripting for automation and API integration.
  • Incremental loading patterns, late-arriving data handling, and reprocessing.
  • Basic AWS S3 and IAM knowledge.
  • Basic understanding of Redshift and its role in architecture.
  • Fluent English.
  • Self-directed and able to navigate unfamiliar codebases without onboarding.
  • Ability to explain production incidents to non-technical stakeholders and provide ETAs.

Responsabilidades

  • Keep Databricks pipelines running for ingestion, transformation, and delivery to downstream consumers.
  • Diagnose and resolve pipeline failures and data quality issues.
  • Reverse-engineer and document existing transformation logic and business rules.
  • Migrate legacy Hive Metastore tables to Unity Catalog.
  • Maintain Iceberg-enabled table sharing between Databricks and Snowflake.
  • Build and test dbt models with incremental materializations and data tests.
  • Develop and maintain Airflow DAGs for orchestration.
  • Validate migrated pipelines against Databricks outputs.
  • Contribute to Snowflake modeling, performance, and cost decisions.
  • Collaborate with client stakeholders on technical topics alongside the team lead.

Conocimientos

SQL
PySpark
Dbt
Airflow
Delta Lake
Unity Catalog
Snowflake
AWS S3
Incremental Loading
Data Quality

Herramientas

Databricks
Snowflake
Airflow
Hive Metastore
Redshift

Descripción del empleo

  • Keep production Databricks pipelines running, including ingestion, transformation, and delivery to downstream consumers
  • Diagnose and resolve pipeline failures and data quality issues
  • Reverse-engineer and document existing transformation logic and business rules
  • Migrate legacy tables from Hive Metastore to Unity Catalog
  • Maintain Iceberg-enabled table sharing between Databricks and Snowflake
  • Build and test dbt models, including incremental materializations and data tests
  • Develop and maintain Airflow DAGs for orchestration
  • Validate migrated pipelines against Databricks outputs
  • Contribute to Snowflake modeling, performance, and cost decisions
  • Work directly with client stakeholders on technical topics alongside the team lead
Requirements
  • 3+ years operating production data pipelines
  • PySpark and SQL — able to read, debug, and modify existing pipelines
  • Delta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioning
  • Databricks Workflows, cluster configuration, job troubleshooting
  • Unity Catalog: catalogs, schemas, grants, lineage, and the metastore model
  • Snowflake warehouses, roles and grants, and general operating model
  • Snowflake query performance and awareness of compute cost behavior
  • dbt models, sources, tests, and incremental materializations
  • dbt project structure and deployment workflow
  • Airflow DAGs, operators, scheduling, and dependency management
  • Airflow retries, backfills, and idempotent task design
  • Strong SQL, including window functions, complex joins, and reading transformation logic
  • Python for scripting, automation, and API integration
  • Incremental loading patterns, idempotency, late-arriving data, and reprocessing
  • AWS S3 and IAM basics
  • Basic working knowledge of Redshift and its role in wider architecture
  • Fluent English
  • Self-directed and able to progress on an unfamiliar codebase without structured onboarding
  • Able to explain production incidents to non-technical stakeholders and provide realistic ETAs
Core Competencies

Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes. Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization.

Highest-signal resume keywords
  • Databricks Pipeline Management
  • SQL Proficiency
  • PySpark Development
  • Airflow DAG Development
  • Dbt Model Building
Hard Skills
  • SQL
  • PySpark
  • Dbt
  • Airflow
  • Delta Lake
  • Unity Catalog
  • Snowflake
  • AWS S3
  • Incremental Loading Patterns
  • Data Quality Diagnosis
Soft Skills
  • Self-Directed
  • Effective Communication
  • Stakeholder Engagement
Industry Keywords
  • Data Pipeline
  • Data Transformation
  • Data Quality
  • Data Modeling
  • Orchestration
Tools & Technologies
  • Databricks
  • Snowflake
  • Airflow
  • Hive Metastore
  • Redshift
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Databricks Data Engineer — Pipelines, dbt & Snowflake
Databricks Data Engineer — Pipelines, dbt & Snowflake

Jobtailor • Barcelona

Presencial
EUR 70.000 - 100.000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

Intetics • España

Presencial
EUR 60.000 - 90.000
1104 | Senior Data Engineer (Databricks)
1104 | Senior Data Engineer (Databricks)

Intetics-Inc • España

Presencial
EUR 60.000 - 90.000
Data Engineer – Snowflake + dbt
Data Engineer – Snowflake + dbt

office people Holding • España

Presencial
EUR 45.000 - 65.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • Zaragoza

Presencial
EUR 70.000 - 100.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • Tarragona

Presencial
EUR 70.000 - 95.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • Badajoz

Presencial
EUR 60.000 - 90.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • Almería

Presencial
EUR 60.000 - 80.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • España

Presencial
EUR 70.000 - 110.000
Data Engineer (Databricks)
Data Engineer (Databricks)

Capitole • Córdoba

Presencial
EUR 70.000 - 100.000