Remote Databricks Data Engineer: Lakehouse & AI Pipelines

AgileEngine

Región Centro

Presencial

MXN 600.000 - 900.000

Jornada completa

hace 25 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Growth without limits
Competitive compensation
Flexibility: 100% remote
Meaningful, modern projects
Collaborative culture
Well-being & support

Descripción de la vacante

AgileEngine is seeking a Middle Data Engineer to modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will design and operate batch and streaming pipelines with PySpark and Delta Lake, applying a medallion architecture across bronze, silver, and gold layers.

You will collaborate with DevOps and analytics engineers, optimize Spark jobs for performance and cost, implement governance with Unity Catalog, and leverage AI tools like Claude and GitHub Copilot to accelerate

Formación

  • 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
  • Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
  • Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
  • Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
  • Strong problem-solving, collaboration, and communication skills.
  • Familiarity with Unity Catalog, data governance, access control, and PII handling.
  • Experience with dbt or an equivalent transformation framework.
  • Familiarity with secure coding standards and industry security best practices.
  • Experience delivering production data platforms at scale.
  • Upper-intermediate English level.

Responsabilidades

  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
  • Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
  • Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
  • Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
  • Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
  • Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.

Conocimientos

Databricks
Spark PySpark
SQL
Python
Data modeling
Unity Catalog
dbt
Structured Streaming
Airflow
Azure Data Factory
Security best practices
PII handling
English (Upper‑intermediate)

Herramientas

Databricks
Terraform
PostgreSQL
Dynatrace
CloudWatch
Databricks Workflows
dbt

Descripción del empleo

AgileEngine is seeking a Middle Data Engineer to modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will design and operate batch and streaming pipelines with PySpark and Delta Lake, applying a medallion architecture across bronze, silver, and gold layers.

You will collaborate with DevOps and analytics engineers, optimize Spark jobs for performance and cost, implement governance with Unity Catalog, and leverage AI tools like Claude and GitHub Copilot to accelerate

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Databricks Data Engineer — Remote Lakehouse & AI Pipelines
Databricks Data Engineer — Remote Lakehouse & AI Pipelines

AgileEngine, LLC. • Rosarito

Presencial
MXN 700.000 - 1.000.000
Growth without limits
Competitive compensation
Remote with flexible hours
+3
Databricks Data Engineer — Remote Lakehouse & AI-Powered Pipelines
Databricks Data Engineer — Remote Lakehouse & AI-Powered Pipelines

AgileEngine • Rosarito

Presencial
MXN 2.036.000 - 2.714.000
Growth without limits
Competitive compensation
100% remote with flexible hours
+3
Databricks Data Engineer — Lakehouse Pipelines & AI Tools
Databricks Data Engineer — Lakehouse Pipelines & AI Tools

AgileEngine • Almoloya de Juárez

Presencial
MXN 480.000 - 720.000
Growth budget
Competitive pay
Remote work
+3
Remote Data Engineer — Lakehouse, PySpark & AI-Accelerated Pipelines
Remote Data Engineer — Lakehouse, PySpark & AI-Accelerated Pipelines

AgileEngine, LLC. • Región Centro

Presencial
MXN 420.000 - 820.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Data Engineer - Databricks Lakehouse & Streaming Pipelines (Remote)
Data Engineer - Databricks Lakehouse & Streaming Pipelines (Remote)

AgileEngine, LLC. • León

Presencial
MXN 1.529.000 - 2.209.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
Remote Senior Databricks Data Engineer — Lakehouse & AI
Remote Senior Databricks Data Engineer — Lakehouse & AI

AgileEngine, LLC. • Rosarito

Presencial
MXN 1.869.000 - 2.718.000
Growth without limits
Competitive compensation
Remote work with flexible hours
+3
Data Engineer - Databricks Lakehouse & Streaming (Remote)
Data Engineer - Databricks Lakehouse & Streaming (Remote)

AgileEngine • Ciudad de México

Presencial
MXN 600.000 - 850.000
Growth opportunities
Competitive pay
Remote work
+3
Senior Databricks Data Engineer - Lakehouse (Remote)
Senior Databricks Data Engineer - Lakehouse (Remote)

AgileEngine • Puebla de Zaragoza

Presencial
MXN 900.000 - 1.200.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Senior Databricks Data Engineer - Remote Lakehouse Excellence
Senior Databricks Data Engineer - Remote Lakehouse Excellence

AgileEngine, LLC. • Puebla de Zaragoza

Presencial
MXN 900.000 - 1.600.000
Growth without limits
Competitive compensation
Remote work with flexible hours
+3
Senior Databricks Data Engineer — Remote Lakehouse Leader
Senior Databricks Data Engineer — Remote Lakehouse Leader

AgileEngine • Ciudad de México

Presencial
MXN 900.000 - 1.200.000
Growth opportunities
Competitive compensation
Remote-friendly
+3