Middle Databricks Data Engineer ID86295

AgileEngine

Bogotá ciudad

Presencial

COP 90.000.000 - 150.000.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Flexible remote work
Annual learning budget
Growth opportunities
Competitive compensation
Supportive culture
Remote collaboration with global teams

Descripción de la vacante

AgileEngine is seeking a Middle Data Engineer to modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. Build batch and streaming data pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers.

Tools like Claude and Copilot speed up development. You will migrate legacy ETL workloads, optimize Spark jobs, and implement data quality and governance with Unity Catalog, while collaborating with DevOps and analytics teams.

Formación

  • 3+ years of professional data engineering experience.
  • Strong hands-on experience with Databricks, Spark (PySpark), and Delta Lake.
  • Advanced SQL and Python with data modeling skills.
  • Experience with streaming ingestion (Structured Streaming, Auto Loader, Kafka, Event Hubs).
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations or ETL modernization.
  • Familiarity with Unity Catalog, data governance, and PII handling.
  • Experience with dbt or equivalent transformation framework.
  • Upper-intermediate English level.

Responsabilidades

  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion architecture serving analytics, reporting, and ML consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal disruption.
  • Use Claude or Copilot to accelerate development and generate tests and docs.
  • Write clean, well-tested Python and SQL with code reviews and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost; manage partitioning and clustering.
  • Implement data quality, lineage, governance via Unity Catalog and automated checks.
  • Debug and resolve production data pipeline incidents; collaborate on observability and security.

Conocimientos

Apache Spark
PySpark
Delta Lake
SQL
Python
Structured Streaming
Databricks Workflows
Data modeling
Data warehouse migration

Herramientas

Databricks
Airflow
Azure Data Factory
Unity Catalog
dbt

Descripción del empleo

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE

We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers. This role also uses AI tools like Claude and GitHub Copilot to speed up development.

WHAT YOU WILL DO
  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
  • Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
  • Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
  • Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
  • Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
  • Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
MUST HAVES
  • 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
  • Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
  • Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
  • Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
  • Strong problem-solving, collaboration, and communication skills.
  • Familiarity with Unity Catalog, data governance, access control, and PII handling.
  • Experience with dbt or an equivalent transformation framework.
  • Familiarity with secure coding standards and industry security best practices.
  • Experience delivering production data platforms at scale.
  • Upper-intermediate English level.
NICE TO HAVES
  • Experience with Infrastructure as Code (IaC) using Terraform and CI/CD using Azure DevOps.
  • Experience working with relational databases (specifically PostgreSQL) and data persistence concepts.
  • Familiarity with logging and monitoring tools (e.g., Dynatrace, CloudWatch, Databricks system tables).
  • Experience working in Agile or team-based development environments preferred.
PERKS AND BENEFITS
  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support: access local well-being programs and people-focused support tailored to your location
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Databricks Lakehouse Data Engineer — Remote
Databricks Lakehouse Data Engineer — Remote

AgileEngine, LLC. • Centrosur

Presencial
COP 6.000.000 - 9.000.000
Remote work
Flexible hours
Learning budget
+2
Remote Data Engineer, Databricks Lakehouse
Remote Data Engineer, Databricks Lakehouse

AgileEngine • Perímetro Urbano Barranquilla

Presencial
COP 281.461.000 - 406.555.000
Growth opportunities
Competitive compensation
Remote work
+3
Remote Data Engineer - Databricks Lakehouse & PySpark
Remote Data Engineer - Databricks Lakehouse & PySpark

AgileEngine • Centrosur

Presencial
COP 100.440.000 - 156.240.000
Growth without limits
Competitive compensation
Flexibility
+3
Data Engineer — Databricks Lakehouse (Remote)
Data Engineer — Databricks Lakehouse (Remote)

AgileEngine • Cartagena de Indias

Presencial
COP 205.918.000 - 300.957.000
Growth opportunities
Competitive compensation
100% remote
+5
Remote Data Engineer: Databricks Lakehouse & Pipelines
Remote Data Engineer: Databricks Lakehouse & Pipelines

AgileEngine • Metropolitana

Presencial
COP 283.751.000 - 409.862.000
Growth without limits
Competitive compensation
Flexibility
+3
Remote Databricks Data Engineer — Lakehouse & AI Pipelines
Remote Databricks Data Engineer — Lakehouse & AI Pipelines

AgileEngine • Bogotá ciudad

Presencial
COP 90.000.000 - 150.000.000
Flexible remote work
Annual learning budget
Growth opportunities
+3
Remote Data Engineer - Databricks Lakehouse & AI Pipelines
Remote Data Engineer - Databricks Lakehouse & AI Pipelines

AgileEngine, LLC. • Bogotá

Presencial
COP 90.000.000 - 150.000.000
Growth budget
Competitive compensation
Remote work
+3
Data Engineer - Remote, Databricks Lakehouse & Pipelines
Data Engineer - Remote, Databricks Lakehouse & Pipelines

AgileEngine, LLC. • Metropolitana

Presencial
COP 90.000.000 - 140.000.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
1104 | Senior Data (Databricks) Engineer
1104 | Senior Data (Databricks) Engineer

Intetics • Colombia

Presencial
COP 108.000.000 - 180.000.000
Senior Databricks Data Engineer | Remote & Growth-Focused
Senior Databricks Data Engineer | Remote & Growth-Focused

AgileEngine, LLC. • Medellín

Presencial
COP 130.000.000 - 170.000.000
Growth without limits
Competitive compensation
Remote work with flexible hours
+3