Transforma esta oferta en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.
Talent To Hire Inc. in Madrid is seeking a hands-on Senior Databricks Data Engineer for a hybrid, six‑month contract with extensions. You will design, build and optimize scalable data pipelines using Databricks, PySpark, Python and SQL.
The role requires working with large datasets, integrating APIs, databases and files, and implementing Bronze, Silver and Gold medallion layers. You will stay hands-on with Databricks/Spark and collaborate with Data Architects and business stakeholders.
Location: Madrid, Spain – Hybrid
Work Hours: Swiss/Spanish business hours
Experience: 6–10+ years
Engagement: Contract ( 6 months with extensions)
Start: Immediate / ASAP
We are looking for a highly hands‑on Senior Databricks Data Engineer to design, build, and optimize scalable end‑to‑end data pipelines.
This role is best suited to an engineer who is comfortable writing PySpark/Python code independently, working with large datasets, integrating multiple data sources, and taking data from initial ingestion through Bronze, Silver, and Gold layers using Medallion Architecture.
This is not a coordination‑only or architecture‑only position. We are specifically seeking someone who remains hands‑on with Databricks, Spark, PySpark, Python and SQL.
Design, develop, and maintain scalable Databricks‑based data pipelines.
Build end‑to‑end data engineering solutions from source ingestion through Gold‑layer datasets.
Develop high‑performance data processing workflows using PySpark, Python, Spark and SQL.
Integrate data from multiple sources, including APIs, databases, files, cloud storage and external platforms.
Design and optimize Databricks architectures for data ingestion, transformation, processing and storage.
Work with large‑scale datasets and distributed data processing environments.
Perform Spark/Databricks performance tuning to improve processing speed, scalability and cost efficiency.
Build robust ETL/ELT workflows with appropriate data quality, monitoring and governance controls.
Optimize Databricks workflows, jobs and compute resources.
Troubleshoot pipeline performance, reliability and data‑quality issues.
Collaborate with Data Architects, Data Engineers and business stakeholders to translate requirements into production‑ready data solutions.
Contribute to engineering standards and Databricks best practices.
you demonstrate strong production‑level experience with:
Databricks
Python
Advanced SQL
Medallion Architecture – Bronze, Silver and Gold
Data ingestion from APIs, databases, files and multiple source systems
Large‑scale/distributed data processing
Data transformation and aggregation
Databricks/Spark performance tuning and optimization
Data quality and pipeline monitoring
Cloud‑based data engineering environments
Experience with some of the following would be advantageous:
AWS: S3, Glue ETL, Lambda, Step Functions, ECS, CloudWatch
Snowflake
DBT
Terraform
BigQuery
Kafka
Git / Jenkins / CI/CD
Data governance
Cost monitoring and cloud optimization
Infrastructure as Code
Databricks Certified Data Engineer Associate or similar Databricks certification is considered an asset.
You are a strong fit if you have personally designed and coded production data pipelines rather than primarily managing other engineers.
You should be able to clearly explain a recent project where you:
and describe the PySpark/Python code, transformations, architecture decisions, performance improvements and data‑quality controls you personally implemented.
Shortlisted candidates should be prepared to discuss:
1. Databricks / Medallion Architecture:
Walk us through a production pipeline you personally built from source ingestion through Bronze, Silver and Gold. What did you personally code?
2. PySpark:
Describe a PySpark pipeline you developed for a large dataset. What transformations did you implement, and how did you optimize its performance?
3. Performance:
A Databricks/Spark job that previously completed in 20 minutes now takes 90 minutes. How would you diagnose and optimize it?
4. Data Ingestion:
How have you ingested data from APIs, relational databases, files or cloud storage into Databricks?
5. Data Quality:
How do you implement data‑quality validation, error handling, monitoring and recovery within a production data pipeline?
6. Optimization:
Give an example where you reduced Databricks/cloud processing costs or significantly improved pipeline performance.
Important: We are prioritizing hands‑on engineers, not candidates whose recent experience is primarily management, coordination or high‑level architecture.