Ingeniero de Datos

PROGRAMMING.COM

Azcapotzalco

Presencial

MXN 520.000 - 760.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

PROGRAMMING.COM is seeking a Data Engineer to build and operate robust data pipelines using Databricks, Airflow, or Dagster. You will develop efficient ETL/ELT workflows in Python and SQL for batch and streaming workloads, and collaborate with ML teams to prepare high‑quality datasets for training and production use.

You will model and maintain Delta/Parquet/Iceberg data assets, implement orchestration and monitoring, and ensure data quality with validation and audit logging.

Formación

  • 3–6 years of experience as a Data Engineer or ETL Developer in a production environment.
  • Proficiency in Python and SQL; strong familiarity with Databricks, Spark, or equivalent big-data frameworks.
  • Experience with workflow orchestration tools such as Airflow, Dagster, Luigi or Prefect.
  • Deep understanding of data modeling, data warehousing, and distributed data processing.
  • Knowledge of modern data lakehouse architectures (Delta, Parquet, Iceberg).
  • Familiarity with CI/CD, GitHub Actions, and data pipeline testing frameworks.
  • Comfort working in a cross-functional environment with ML, product, and analytics teams.

Responsabilidades

  • Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Dagster.
  • Develop efficient ETL/ELT workflows in Python and SQL to support both batch and streaming workloads.
  • Collaborate with ML and AI teams to deliver high-quality datasets for training, evaluation, and production features.
  • Model and maintain structured data assets (Delta, Parquet, Iceberg) for reliability, versioning, and lineage tracking.
  • Implement orchestration and monitoring — schedule jobs, track dependencies, and automate recovery from failures.
  • Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging.
  • Contribute to data platform evolution — evaluate tools, standardize best practices, and improve developer experience.
  • Support performance and cost optimization across compute, storage, and orchestration systems.

Conocimientos

Python
SQL
Data Modeling
Big Data
CI/CD
GitHub Actions

Herramientas

Databricks
Airflow
Dagster
Spark

Descripción del empleo

Role: Data Engineer

Location: Remote (Requires occasional travel)

Responsibilities
  • Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Dagster.
  • Develop efficient ETL/ELT workflows in Python and SQL to support both batch and streaming workloads.
  • Collaborate with ML and AI teams to deliver high-quality datasets for training, evaluation, and production features.
  • Model and maintain structured data assets (Delta, Parquet, Iceberg) for reliability, versioning, and lineage tracking.
  • Implement orchestration and monitoring — schedule jobs, track dependencies, and automate recovery from failures.
  • Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging.
  • Contribute to data platform evolution — evaluate tools, standardize best practices, and improve developer experience.
  • Support performance and cost optimization across compute, storage, and orchestration systems.
Qualifications
  • 3–6 years of experience as a Data Engineer or ETL Developer in a production environment.
  • Proficiency in Python and SQL; strong familiarity with Databricks, Spark, or equivalent big-data frameworks.
  • Experience with workflow orchestration tools such as Airflow, Dagster, Luigi or Prefect.
  • Deep understanding of data modeling, data warehousing, and distributed data processing.
  • Knowledge of modern data lakehouse architectures (Delta, Parquet, Iceberg).
  • Familiarity with CI/CD, GitHub Actions, and data pipeline testing frameworks.
  • Comfort working in a cross-functional environment with ML, product, and analytics teams.
Nice to Have
  • Experience with sports, telemetry, or sensor data pipelines.
  • Familiarity with streaming frameworks (Kafka, Spark Structured Streaming, Flink).
  • General knowledge of American football, the NFL, and college football
  • Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog,OpenLineage).
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and containerization (Docker, Kubernetes).
  • Exposure to best practices in machine-learning model management and MLOps
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Data Engineer
Data Engineer

Pyramid Consulting, Inc • Región Centro

Presencial
MXN 900.000 - 1.500.000
Solution Data Engineer Sr.
Solution Data Engineer Sr.

Turtle Trax S.A. • Región Centro

Presencial
MXN 1.224.000 - 1.575.000
Data Engineer Sr
Data Engineer Sr

Turtle Trax S.A. • Región Centro

Híbrido
MXN 800.000 - 1.100.000
Data Engineer
Data Engineer

Corus Consulting an Alten Company • México

A distancia
MXN 600.000 - 900.000
Data Engineer Lead
Data Engineer Lead

Turtle Trax S.A. • Región Centro

Híbrido
MXN 1.200.000 - 1.800.000
Senior AI Data Engineer
Senior AI Data Engineer

Strategic Systems International • Ciudad de México

Presencial
MXN 900.000 - 1.500.000
Senior Data Engineer
Senior Data Engineer

Nimble Gravity • Región Centro

Presencial
MXN 600.000 - 800.000
Data Engineer - Databricks
Data Engineer - Databricks

Zurich Insurance • Región Centro

Presencial
MXN 860.000 - 1.205.000
Remote Data Engineer — Databricks Lakehouse & AI Tools
Remote Data Engineer — Databricks Lakehouse & AI Tools

AgileEngine • Puebla de Zaragoza

Presencial
MXN 420.000 - 640.000
Growth without limits
Competitive compensation
Flexibility
+3
Remote Data Engineer — Scalable Pipelines & Lakehouse
Remote Data Engineer — Scalable Pipelines & Lakehouse

PROGRAMMING.COM • Azcapotzalco

Presencial
MXN 520.000 - 760.000