Middle Databricks Data Engineer ID86295

AgileEngine

Ciudad de México

Presencial

MXN 600.000 - 850.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Growth opportunities
Competitive pay
Remote work
Modern projects
Collaborative culture
Well-being programs

Descripción de la vacante

AgileEngine is seeking a Middle Data Engineer to modernize a data warehouse to a Databricks Lakehouse. You will design batch and streaming pipelines with PySpark and Delta Lake across bronze, silver, and gold layers, leveraging AI tools to accelerate development.

The role emphasizes data governance, production-grade pipelines, and collaboration with DevOps and analytics teams. Remote work options are available.

Formación

  • 3+ years of data engineering with Spark and cloud architectures.
  • Strong hands-on with Databricks, Spark (PySpark) and Delta Lake.
  • Advanced SQL and Python with Lakehouse data modeling.
  • Streaming ingestion using Structured Streaming, Auto Loader, Kafka or Event Hubs.
  • Workflow orchestration with Databricks Workflows, Airflow or Azure Data Factory.
  • Experience migrating legacy ETL and modernizing data warehouses.
  • Strong problem-solving, collaboration, and communication.
  • Familiarity with Unity Catalog, data governance and PII handling.
  • Experience with dbt or comparable transformation framework.
  • Secure coding practices and security standards.
  • Experience delivering production data platforms at scale.
  • Upper-intermediate English level.

Responsabilidades

  • Design, build, and operate batch and streaming pipelines on Databricks.
  • Model and maintain a medallion architecture for analytics and ML.
  • Migrate workloads to Lakehouse with data parity and minimal disruption.
  • Utilize Claude or Copilot to accelerate development and testing.
  • Write clean Python and SQL with thorough code reviews.
  • Optimize Spark jobs and Delta tables for performance and cost.
  • Implement data quality, lineage, and governance using Unity Catalog.
  • Debug and resolve pipeline failures and data defects.
  • Collaborate with DevOps, platform, and analytics teams on observability and security.

Conocimientos

Databricks
PySpark
Spark
Delta Lake
SQL
Python
Structured Streaming
Kafka
Event Hubs
Airflow
Databricks Workflows
Azure Data Factory
Unity Catalog
dbt
Security practices
English Proficiency
Production data platforms

Herramientas

Terraform
CloudWatch
PostgreSQL

Descripción del empleo

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you’re looking for a place to grow, make an impact, and work with people who care, we’d love to meet you!

ABOUT THE ROLE

We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers. This role also uses AI tools like Claude and Github Copilot to speed up development.

WHAT YOU WILL DO
  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
  • Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
  • Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
  • Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
  • Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
  • Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
MUST HAVES
  • 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
  • Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
  • Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
  • Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
  • Strong problem-solving, collaboration, and communication skills.
  • Familiarity with Unity Catalog, data governance, access control, and PII handling.
  • Experience with dbt or an equivalent transformation framework.
  • Familiarity with secure coding standards and industry security best practices.
  • Experience delivering production data platforms at scale.
  • Upper-intermediate English level.
NICE TO HAVES
  • Experience with Infrastructure as Code (IaC) using Terraform and CI/CD using Azure DevOps.
  • Experience working with relational databases (specifically PostgreSQL) and data persistence concepts.
  • Familiarity with logging and monitoring tools (e.g., Dynatrace, CloudWatch, Databricks system tables).
  • Experience working in Agile or team-based development environments preferred.
PERKS AND BENEFITS
  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support: access local well-being programs and people-focused support tailored to your location
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Data Engineer — Databricks Lakehouse & AI Tools
Remote Data Engineer — Databricks Lakehouse & AI Tools

AgileEngine • Puebla de Zaragoza

Presencial
MXN 420.000 - 640.000
Growth without limits
Competitive compensation
Flexibility
+3
Senior Databricks Data Engineer ID86297
Senior Databricks Data Engineer ID86297

AgileEngine, LLC. • Región Centro

A distancia
MXN 600.000 - 1.200.000
Growth without limits
Competitive compensation
Flexibility
+3
Data Engineer - Databricks Lakehouse & Streaming Pipelines (Remote)
Data Engineer - Databricks Lakehouse & Streaming Pipelines (Remote)

AgileEngine, LLC. • León

Presencial
MXN 1.529.000 - 2.209.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
Databricks Data Engineer — Remote Lakehouse & AI Pipelines
Databricks Data Engineer — Remote Lakehouse & AI Pipelines

AgileEngine, LLC. • Rosarito

Presencial
MXN 700.000 - 1.000.000
Growth without limits
Competitive compensation
Remote with flexible hours
+3
Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine, LLC. • Región Centro

Presencial
MXN 900.000 - 1.200.000
Growth opportunities
Competitive pay
Remote work 100%
+3
Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine, LLC. • Santiago de Querétaro

Presencial
MXN 2.042.000 - 2.893.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Senior Databricks Data Engineer — Remote, Flexible Hours
Senior Databricks Data Engineer — Remote, Flexible Hours

AgileEngine, LLC. • Monterrey

Presencial
MXN 700.000 - 1.100.000
Growth
Compensation
Remote work
+3
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Rosarito

Híbrido
MXN 1.612.000 - 2.151.000
Mentorship programs
Competitive compensation
Flexible schedule
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Santiago de Querétaro

Híbrido
MXN 2.042.000 - 3.063.000
Professional growth with mentorship and TechTalks
Competitive compensation with various budgets
Engaging projects with Fortune 500 clients
+1
Remote Senior Databricks Data Engineer — Lakehouse & AI
Remote Senior Databricks Data Engineer — Lakehouse & AI

AgileEngine, LLC. • Rosarito

Presencial
MXN 1.869.000 - 2.718.000
Growth without limits
Competitive compensation
Remote work with flexible hours
+3