Remote Data Engineer: Spark, Databricks & ETL Pipelines

Capgemini Engineering

Aguascalientes

Presencial

MXN 420.000 - 660.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Capgemini Engineering is seeking a MID -Data Engineer- for a REMOTE role. You will design, build, and optimize data pipelines using Databricks and Spark, integrate data from databases, APIs, and streaming sources, and work closely with data scientists to enable advanced analytics.

You will implement data governance, ensure data quality, and support CI/CD for data workflows across cloud platforms such as Azure, AWS, and GCP.

Formación

  • Hands-on experience with data engineering or big data development.
  • Strong knowledge of SQL / Oracle / SQL Server.
  • Strong knowledge of Java / JavaScript.
  • Strong knowledge of Web-based applications and APIs (REST/SOAP).
  • Strong experience with Databricks Platform.
  • Strong experience with Apache Spark (PySpark/Scala).
  • Strong experience with SQL & Python.
  • Experience with Delta Lake and data lake architecture.
  • Hands-on experience with cloud platforms (Azure, AWS, or GCP).
  • Familiarity with data orchestration tools (Azure Data Factory, Airflow, etc.).
  • Knowledge of data warehousing concepts (Star schema, Snowflake schema).
  • Experience with version control (Git) and CI/CD pipelines.
  • Strong understanding of data pipeline optimization and performance tuning.

Responsabilidades

  • Design, develop, and support data replication and integration solutions using HVR.
  • Design, build, and maintain scalable data pipelines using Databricks (Spark, Delta Lake).
  • Develop and optimize ETL/ELT processes for structured and unstructured data.
  • Work with large datasets to ensure data quality, integrity, and performance optimization.
  • Implement data models and transformations for analytics and reporting.
  • Collaborate with data scientists and analysts to enable advanced analytics and ML workloads.
  • Integrate data from multiple sources including databases, APIs, and streaming systems.
  • Optimize Spark jobs for performance tuning and cost efficiency.
  • Implement data governance, security, and access controls.
  • Monitor and troubleshoot data pipelines and production issues.
  • Support CI/CD pipelines and DevOps best practices for data engineering workflows.
  • Support data migration and modernization initiatives.
  • Ensure data quality, governance, security, and compliance standards.
  • Create operational documentation, runbooks, and support procedures.
  • Participate in production support, issue resolution, and performance tuning activities.

Conocimientos

Big data
SQL
JavaScript
Python
Data governance

Educación

Bachelor’s degree in Computer Science or related field

Herramientas

Databricks
Apache Spark
Delta Lake
Azure
AWS
GCP
REST APIs

Descripción del empleo

Capgemini Engineering is seeking a MID -Data Engineer- for a REMOTE role. You will design, build, and optimize data pipelines using Databricks and Spark, integrate data from databases, APIs, and streaming sources, and work closely with data scientists to enable advanced analytics.

You will implement data governance, ensure data quality, and support CI/CD for data workflows across cloud platforms such as Azure, AWS, and GCP.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Data Engineer – Databricks & Cloud Pipelines
Remote Data Engineer – Databricks & Cloud Pipelines

Capgemini • Aguascalientes

Presencial
MXN 520.000 - 760.000
Remote Senior Data Engineer - Lakehouse & Spark
Remote Senior Data Engineer - Lakehouse & Spark

Capgemini Engineering • Aguascalientes

Presencial
MXN 600.000 - 900.000
Flexible work location
Learning opportunities
Remote Data Engineer: Azure Synapse & PySpark
Remote Data Engineer: Azure Synapse & PySpark

Bluelight Consulting LLC • Chihuahua

A distancia
MXN 600.000 - 1.100.000
Senior Data Engineer - Databricks & Azure (Remote LATAM)
Senior Data Engineer - Databricks & Azure (Remote LATAM)

ITH. • Ciudad de México

Presencial
MXN 1.393.000 - 2.091.000
Fully remote opportunity
Long-term collaboration potential
Competitive compensation in USD
Senior Databricks Engineer (PySpark & Delta Lake)
Senior Databricks Engineer (PySpark & Delta Lake)

Apex Systems • Región Centro

Presencial
MXN 1.099.000 - 1.222.000
Remote Data Engineer - Spark, Delta Lake, Python (Flexible)
Remote Data Engineer - Spark, Delta Lake, Python (Flexible)

BairesDev • Jalisco

A distancia
MXN 200.000 - 400.000
Junior Cloud Data Engineer — Build Data Pipelines
Junior Cloud Data Engineer — Build Data Pipelines

Capgemini • Aguascalientes

Presencial
MXN 240.000 - 360.000
Senior Data Engineer — Databricks & Azure Data Pipelines
Senior Data Engineer — Databricks & Azure Data Pipelines

PwC Acceleration Centers • Estado de México

Presencial
MXN 480.000 - 720.000
Senior Data Engineer - Databricks on AWS (Remote)
Senior Data Engineer - Databricks on AWS (Remote)

Lumenalta • Estado de México

Presencial
MXN 1.426.000 - 2.140.000
Remote Data Engineer – Azure PySpark & ETL
Remote Data Engineer – Azure PySpark & ETL

Bluelight Consulting LLC • Mexicali

A distancia
MXN 400.000 - 700.000