Data Engineer, Databricks

Jobtailor

Lisboa

Presencial

EUR 50 000 - 70 000

Tempo integral

Há 5 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

Jobtailor in Lisbon is seeking a data engineer to build and maintain scalable data pipelines on the Databricks Lakehouse Platform using Spark and PySpark. You will ingest data from databases, files and APIs, model bronze/silver/gold transforms, and ensure data quality and governance.

You will optimize performance and costs, apply security and Unity Catalog governance, contribute to CI/CD for data assets, and collaborate with data scientists and analysts to deliver reliable datasets for analytics

Qualificações

  • Bachelor's or Master's degree in CS, IS, Engineering, or related field.
  • 2+ years of professional experience in data engineering or software engineering with a strong data component.
  • Solid Python skills in production environments.
  • Hands-on experience with Spark/PySpark for large-scale batch or streaming data processing.

Responsabilidades

  • Build and maintain scalable data processing pipelines on the Databricks Lakehouse Platform using Apache Spark and PySpark.
  • Ingest and integrate datasets from databases, file systems, and APIs into the Lakehouse using Auto Loader, CDC, and batch or streaming ingestion patterns.
  • Model and implement bronze, silver, and gold transformation layers using Delta Lake and Databricks SQL.
  • Ensure data integrity, consistency, and quality through validation and monitoring using Delta Live Tables expectations and Lakehouse monitoring capabilities.
  • Tune pipelines for performance and cost through Spark job optimization, Delta table maintenance, and cluster and compute configuration.
  • Apply data security, access control, and privacy practices using Unity Catalog for governance, permissions, and lineage.
  • Contribute to CI/CD and deployment automation for data assets.
  • Support migration initiatives from on-premise Cloudera environments to the Databricks Lakehouse.
  • Collaborate with data scientists and analysts to deliver reliable datasets for analytics and machine learning workflows.
  • Occasionally support lightweight Python services and APIs exposing data to downstream consumers.

Conhecimentos

Python
SQL
Apache Spark
PySpark
Data Modeling
Delta Lake
Delta Live Tables
Structured Streaming
Data Ingestion
CI/CD
Testing
Version Control
Collaboration
Communication

Formação académica

Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field

Ferramentas

Databricks
Airflow
Kafka
Terraform

Descrição da oferta de emprego

  • Build and maintain scalable data processing pipelines and workflows on the Databricks Lakehouse Platform using Apache Spark and PySpark
  • Ingest and integrate datasets from databases, file systems, and APIs into the Lakehouse using Auto Loader, CDC, and batch or streaming ingestion patterns
  • Model and implement bronze, silver, and gold transformation layers using Delta Lake and Databricks SQL
  • Ensure data integrity, consistency, and quality through validation and monitoring using Delta Live Tables expectations and Lakehouse monitoring capabilities
  • Tune pipelines for performance and cost through Spark job optimization, Delta table maintenance, and cluster and compute configuration
  • Apply data security, access control, and privacy practices using Unity Catalog for governance, permissions, and lineage
  • Contribute to CI/CD and deployment automation for data assets
  • Support migration initiatives from on-premise Cloudera environments to the Databricks Lakehouse
  • Collaborate with data scientists and analysts to deliver reliable datasets for analytics and machine learning workflows
  • Occasionally support lightweight Python services and APIs exposing data to downstream consumers
Requirements
  • Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field
  • 2+ years of professional experience in data engineering or software engineering with a strong data component
  • Solid Python skills in production environments
  • Good engineering practices, including version control, testing, and code review
  • Hands-on experience with Spark/PySpark for large-scale batch or streaming data processing
  • Working experience with Databricks or a comparable Lakehouse/cloud data platform
  • Experience with one major cloud provider: Azure, AWS, or GCP
  • Strong SQL skills
  • Solid data modelling fundamentals across relational and non-relational stores
  • Fluency in English, written and spoken
  • Databricks certifications are nice to have
  • Experience with Delta Live Tables, Structured Streaming, or Unity Catalog is nice to have
  • Familiarity with Databricks Workflows, Airflow, Kafka, Terraform, Databricks Asset Bundles, CI/CD tooling, Photon, Databricks SQL Warehouses, or serverless compute is nice to have
Core Competencies

Demonstrates expertise in building and maintaining scalable data processing pipelines on the Databricks Lakehouse Platform, utilizing Apache Spark and PySpark. Proficient in data ingestion, transformation, and ensuring data integrity while applying best practices in data security and governance.

Highest-signal resume keywords
  • Apache Spark
  • PySpark
  • Databricks Lakehouse
  • SQL
  • Data Engineering
Hard Skills
  • Python
  • Data Modeling
  • Delta Lake
  • Delta Live Tables
  • Structured Streaming
  • CI/CD
  • Version Control
  • Testing
  • Code Review
  • Data Ingestion
Soft Skills
  • Collaboration
  • Communication
Certifications & Qualifications
  • Databricks Certifications
Industry Keywords
  • Cloud Data Platform
  • Data Processing Pipelines
  • Data Integrity
  • Data Quality
  • Data Security
Tools & Technologies
  • Auto Loader
  • Unity Catalog
  • Airflow
  • Kafka
  • Terraform
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Data Engineer (Databricks Migration)
Senior Data Engineer (Databricks Migration)

Sigma Software • Lisboa

Presencial
EUR 70 000 - 120 000
Senior Data Engineer (Databricks)
Senior Data Engineer (Databricks)

DataOnline Corp. • Porto

Presencial
EUR 55 000 - 90 000
Senior Data Engineer (Databricks)
Senior Data Engineer (Databricks)

Anova • Porto

Híbrido
EUR 55 000 - 78 000
Data Engineer
Data Engineer

Intellias • Portugal

Presencial
EUR 70 000 - 100 000
Health Insurance
Permanent contract
Office in Porto (modern, well-equipped
+2
Data Engineer
Data Engineer

Xgeeks • Leiria, Lisboa, Viseu

Híbrido
EUR 30 000 - 50 000
Global Projects & Opportunities
Social Events & Team Building
Training & Development
+6
Data Platform Lead
Data Platform Lead

Confidential • Lisboa

Presencial
EUR 70 000 - 90 000
Data Engineer
Data Engineer

Acolad Finland Oy • Lisboa

Presencial
EUR 42 000 - 66 000
Senior Data Engineer
Senior Data Engineer

Dsv Inc • Lisboa

Híbrido
EUR 70 000 - 110 000
Permanent contract
Private medical care
English courses
+4
Senior Data Engineer - Databricks Lakehouse (Remote)
Senior Data Engineer - Databricks Lakehouse (Remote)

AgileEngine • Porto

Presencial
EUR 65 000 - 90 000
Growth without limits
Competitive compensation
Flexibility: fully remote
Senior/Lead Data Engineer - Lisbon
Senior/Lead Data Engineer - Lisbon

Opplane • Lisboa

Presencial
EUR 60 000 - 80 000
Office Snacks and Activities
Collaborative Team Culture