Databricks Data Engineer — Lakehouse, PySpark & AI Tools

AgileEngine, LLC.

Medellín

Presencial

COP 289.119.000 - 417.617.000

Jornada completa

Hace 10 días
Generador de candidaturas

Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Growth without limits
Competitive compensation
Flexibility: 100% remote
Meaningful, modern projects
Collaborative culture
Well-being & support

Descripción de la vacante

AgileEngine is seeking a Middle Data Engineer to modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture for analytics and ML workloads.

You will leverage Claude and GitHub Copilot to accelerate development, write robust Python/SQL, and ensure data quality and governance across pipelines.

Formación

  • 3+ years of professional experience in data engineering with Spark and cloud data architectures.
  • Hands-on experience building data pipelines with Databricks, PySpark, and Delta Lake.
  • Advanced SQL and Python with strong data modeling skills.
  • Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations, ETL modernization, or data hygiene during porting.
  • Strong problem-solving, collaboration, and communication skills.
  • Familiarity with Unity Catalog, data governance, access control, and PII handling.
  • Experience with dbt or equivalent transformation framework.
  • Familiarity with secure coding standards and industry security best practices.

Responsabilidades

  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and ML consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity.
  • Use Claude or Github Copilot to accelerate development, generate code scaffolding, and review tests.
  • Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost with partitioning, clustering, and caching.
  • Implement data quality, lineage, and governance using Unity Catalog and automated checks.
  • Debug, troubleshoot, and resolve pipeline failures and production incidents.
  • Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance.

Conocimientos

Apache Spark (PySpark)
Delta Lake
SQL
Python
Streaming pipelines
Databricks Workflows
Airflow
Azure Data Factory
Unity Catalog
dbt
KPI data modeling

Herramientas

Databricks

Descripción del empleo

AgileEngine is seeking a Middle Data Engineer to modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture for analytics and ML workloads.

You will leverage Claude and GitHub Copilot to accelerate development, write robust Python/SQL, and ensure data quality and governance across pipelines.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Data Engineer - Databricks Lakehouse & PySpark
Remote Data Engineer - Databricks Lakehouse & PySpark

AgileEngine • Centrosur

Presencial
COP 100.440.000 - 156.240.000
Growth without limits
Competitive compensation
Flexibility
+3
Data Engineer - Databricks Lakehouse & PySpark
Data Engineer - Databricks Lakehouse & PySpark

AgileEngine • Capital

Presencial
COP 189.167.000 - 315.278.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Data Engineer - Databricks Lakehouse & Pipelines
Data Engineer - Databricks Lakehouse & Pipelines

AgileEngine, LLC. • Perímetro Urbano Barranquilla

Presencial
COP 289.119.000 - 417.617.000
Growth without limits
Competitive compensation
Remote work with flexible hours
+3
Databricks Lakehouse Data Engineer — Remote
Databricks Lakehouse Data Engineer — Remote

AgileEngine, LLC. • Centrosur

Presencial
COP 6.000.000 - 9.000.000
Remote work
Flexible hours
Learning budget
+2
Data Engineer — Databricks Lakehouse (Remote)
Data Engineer — Databricks Lakehouse (Remote)

AgileEngine • Cartagena de Indias

Presencial
COP 205.918.000 - 300.957.000
Growth opportunities
Competitive compensation
100% remote
+5
Remote Data Engineer: Databricks Lakehouse & Pipelines
Remote Data Engineer: Databricks Lakehouse & Pipelines

AgileEngine • Metropolitana

Presencial
COP 283.751.000 - 409.862.000
Growth without limits
Competitive compensation
Flexibility
+3
Remote Data Engineer — Databricks Lakehouse & Pipelines
Remote Data Engineer — Databricks Lakehouse & Pipelines

AgileEngine • Bogotá

Presencial
COP 189.167.000 - 283.751.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Remote Data Engineer - Databricks Lakehouse & AI Pipelines
Remote Data Engineer - Databricks Lakehouse & AI Pipelines

AgileEngine, LLC. • Bogotá

Presencial
COP 90.000.000 - 150.000.000
Growth budget
Competitive compensation
Remote work
+3
Data Engineer - Databricks Lakehouse & AI-Driven Pipelines
Data Engineer - Databricks Lakehouse & AI-Driven Pipelines

AgileEngine • Pereira

Presencial
COP 84.000.000 - 126.000.000
Growth budget
Competitive pay
Remote work options
+3
Remote Senior Databricks Data Engineer — Lakehouse & AI
Remote Senior Databricks Data Engineer — Lakehouse & AI

AgileEngine, LLC. • Centrosur

Presencial
COP 385.493.000 - 578.239.000
Growth opportunities
Competitive compensation
Remote-friendly
+3