Data Engineer

Capgemini Engineering

Aguascalientes

Presencial

MXN 600.000 - 900.000

Jornada completa

14 días+
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo: un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Flexible work arrangements
Continuous learning opportunities

Descripción de la vacante

Capgemini Engineering is seeking a MID -Data Engineer- to design, build, and optimize scalable data pipelines and analytics solutions on the Databricks platform. You will leverage Python, SQL, and cloud services to transform and load data for analytics across engineering, manufacturing, and supply chain domains.

The role focuses on data replication, ETL/ELT development, governance, and collaboration with data scientists to enable ML workloads in a distributed environment.

Formación

  • Bachelor’s degree in computer science, engineering, or related field.
  • Hands-on experience with data engineering or big data development.
  • Strong knowledge of SQL, Oracle, or SQL Server.
  • Experience with Databricks Platform and Apache Spark (PySpark/Scala).
  • Proficiency in Python and SQL; familiarity with Delta Lake and data lake architecture.
  • Cloud experience (Azure, AWS, or GCP) and data orchestration tools (Azure Data Factory, Airflow).
  • Understanding of data warehousing concepts and version control (Git).

Responsabilidades

  • Design, develop, and support data replication and integration solutions using HVR.
  • Build and maintain scalable data pipelines with Databricks (Spark, Delta Lake).
  • Develop and optimize ETL/ELT processes for structured and unstructured data.
  • Ensure data quality, integrity, and performance across large datasets.
  • Implement data models and transformations for analytics and reporting.
  • Collaborate with data scientists and analysts for advanced analytics and ML workloads.
  • Integrate data from databases, APIs, and streaming systems.
  • Optimize Spark jobs for performance and cost efficiency.
  • Implement data governance, security, and access controls.
  • Monitor and troubleshoot data pipelines and production issues.
  • Support CI/CD pipelines and DevOps practices for data engineering workflows.
  • Assist data migration and modernization initiatives.
  • Ensure data quality, governance, security, and compliance.
  • Create operational documentation and runbooks.
  • Participate in production support and performance tuning activities.

Conocimientos

Databricks Platform
Python
SQL
Spark
Delta Lake
Azure
Power BI
Tableau
Git
CI/CD

Educación

Bachelor’s degree in computer science or engineering

Herramientas

Databricks Platform
Delta Lake
Azure Data Factory
Airflow
Git
CI/CD tooling
Power BI
Tableau

Descripción del empleo

MID -Data Engineer- | REMOTE

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.

We are seeking a skilled Databricks Engineer / Developer to design, develop and optimize scalable data pipelines and analytics solutions using the Databricks platform. The ideal candidate will have strong expertise in Python/Azure/Power BI, cloud data engineering and experience working in distributed data processing environments. This role involves the development and application of engineering practice and knowledge in defining, configuring and deploying industrial digital technologies (including but not limited to PLM and MES) for managing continuity of information across the engineering enterprise, including design, industrialization, manufacturing and supply chain, and for managing the manufacturing data.

YOUR ROLE
  • Design, develop, and support data replication and integration solutions using HVR
  • Design, build, and maintain scalable data pipelines using Databricks (Spark, Delta Lake)
  • Develop and optimize ETL/ELT processes for structured and unstructured data
  • Work with large datasets to ensure data quality, integrity, and performance optimization
  • Implement data models and transformations for analytics and reporting
  • Collaborate with data scientists and analysts to enable advanced analytics and ML workloads
  • Integrate data from multiple sources including databases, APIs, and streaming systems
  • Optimize Spark jobs for performance tuning and cost efficiency
  • Implement data governance, security, and access controls
  • Monitor and troubleshoot data pipelines and production issues
  • Support CI/CD pipelines and DevOps best practices for data engineering workflows
  • Support data migration and modernization initiatives.
  • Ensure data quality, governance, security, and compliance standards.
  • Create operational documentation, runbooks, and support procedures.
  • Participate in production support, issue resolution, and performance tuning activities.
YOUR PROFILE
  • Bachelor’s degree incomputer science, Engineering, or related field
  • Technical Skills:
  • Hands-on experience with data engineering or big data development
  • Strong knowledge of:
  • SQL / Oracle / SQL Server
  • Java / JavaScript (preferred)
  • Web-based applications and APIs (REST/SOAP)
  • Strong experience with:
  • Databricks Platform
  • Apache Spark (PySpark/Scala)
  • SQL & Python
  • Experience with Delta Lake and data lake architecture
  • Hands-on experience with cloud platforms (Azure, AWS, or GCP)
  • Familiarity with data orchestration tools (Azure Data Factory, Airflow, etc.)
  • Knowledge of data warehousing concepts (Star schema, Snowflake schema)
  • Experience with version control (Git) and CI/CD pipelines
  • Strong understanding of data pipeline optimization and performance tuning
Tech Stack:
  • Databricks (Spark, Delta Lake)
  • Python / PySpark / SQL
  • Azure (ADLS, Synapse, ADF) / AWS (S3, Redshift, Glue)
  • Git, CI/CD tools
  • Data visualization tools (Power BI, Tableau)
  • Bachelor’s degree incomputer science, Engineering, or related field
  • Technical Skills:
  • Hands-on experience with data engineering or big data development
  • Strong knowledge of:
  • SQL / Oracle / SQL Server
  • Java / JavaScript (preferred)
  • Web-based applications and APIs (REST/SOAP)
  • Strong experience with:
  • Databricks Platform
  • Apache Spark (PySpark/Scala)
  • SQL & Python
  • Experience with Delta Lake and data lake architecture
  • Hands-on experience with cloud platforms (Azure, AWS, or GCP)
  • Familiarity with data orchestration tools (Azure Data Factory, Airflow, etc.)
  • Knowledge of data warehousing concepts (Star schema, Snowflake schema)
  • Experience with version control (Git) and CI/CD pipelines
  • Strong understanding of data pipeline optimization and performance tuning
Tech Stack:
  • Databricks (Spark, Delta Lake)
  • Python / PySpark / SQL
  • Azure (ADLS, Synapse, ADF) / AWS (S3, Redshift, Glue)
  • Git, CI/CD tools
  • Data visualization tools (Power BI, Tableau)
WHAT YOU’LL LOVE ABOUT WORKING HERE?
  • At Capgemini Engineering, we encourage flexibility in how, when, and where people get their work done, allowing a better work-life balance, and greater empowerment. They partner with their managers to find an arrangement that works best for their role and their circumstances.
  • At Capgemini Engineering, we’re always looking ahead. We’re part of a team that creates opportunities to achieve valuable change. Change that makes a difference. New connections, new technologies, new ways to work. It’s so energizing.
  • At Capgemini Engineering, we make it easy for you to deepen knowledge and learn new skills while you’re still doing the day job.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Senior Data Engineer - Lakehouse & Spark
Remote Senior Data Engineer - Lakehouse & Spark

Capgemini Engineering • Aguascalientes

Presencial
MXN 600.000 - 900.000
Flexible work location
Learning opportunities
Sr Data Engineer
Sr Data Engineer

Capgemini Engineering • Aguascalientes

Presencial
MXN 600.000 - 900.000
Flexible work location
Learning opportunities
Remote Data Engineer – Databricks & Cloud Pipelines
Remote Data Engineer – Databricks & Cloud Pipelines

Capgemini • Aguascalientes

Presencial
MXN 520.000 - 760.000
Remote Databricks Data Engineer | Azure & Power BI
Remote Databricks Data Engineer | Azure & Power BI

Capgemini Engineering • Aguascalientes

A distancia
MXN 600.000 - 900.000
Flexible work arrangements
Continuous learning opportunities
Senior Databricks Engineer ID86295
Senior Databricks Engineer ID86295

AgileEngine, LLC. • Rosarito

Presencial
MXN 2.063.000 - 3.095.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Senior Databricks Engineer ID86295
Senior Databricks Engineer ID86295

AgileEngine, LLC. • Región Centro

Híbrido
MXN 450.000 - 1.400.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Senior Databricks Engineer ID86295
Senior Databricks Engineer ID86295

AgileEngine, LLC. • Santiago de Querétaro

Híbrido
MXN 1.891.000 - 2.751.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Data Engineer
Data Engineer

ITJ • Tijuana

Presencial
MXN 480.000 - 980.000
Health insurance
Dental insurance
Vision insurance
+6
Devops Support
Devops Support

Capgemini • Aguascalientes

Presencial
MXN 520.000 - 760.000
Jr Cloud Data Engineer
Jr Cloud Data Engineer

Capgemini • Aguascalientes

Presencial
MXN 240.000 - 360.000