Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSSTech Group

United Arab Emirates

On-site

AED 250,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GSSTech Group in United Arab Emirates seeks an experienced Senior Data Engineer to design, build and optimize large-scale data pipelines using PySpark, Python, and Cloudera CDP. You will collaborate with data architects, data scientists, and business stakeholders to deliver scalable data solutions powering BI, analytics, and ML initiatives.

The ideal candidate has 8–10+ years in data engineering, strong SQL, data modelling, data quality, and governance capabilities, and hands-on CDP experience.

Qualifications

  • 8–10+ years of experience in Data Engineering.
  • Strong hands-on expertise in Python.
  • Strong hands-on expertise in PySpark.
  • Extensive experience with Cloudera Data Platform (CDP).
  • Strong understanding of distributed data processing frameworks.
  • Experience building enterprise-scale ETL/ELT pipelines.
  • Strong knowledge of Big Data technologies.
  • Experience with Hadoop ecosystem technologies.
  • Strong SQL programming and query optimization skills.
  • Experience with data modelling and data transformation techniques.
  • Knowledge of data quality, validation, and governance principles.
  • Experience working with Git and version control systems.
  • Understanding of CI/CD practices for data engineering.
  • Strong understanding of REST APIs and data integration patterns.
  • Experience working with structured, semi-structured, and unstructured datasets.
  • Nice to Have Experience working with cloud-native data platforms.
  • Exposure to machine learning data pipelines.
  • Exposure to containerization technologies such as Docker and Kubernetes.
  • Experience with DevOps practices for Data Engineering.

Responsibilities

  • Design, develop, and maintain scalable, high-performance data pipelines using PySpark and Python.
  • Build robust batch and distributed data processing solutions using CDP.
  • Develop and optimize ETL/ELT pipelines for structured and unstructured enterprise datasets.
  • Design scalable data ingestion frameworks from multiple enterprise data sources.
  • Ensure high data quality, integrity, governance, and availability across enterprise platforms.
  • Perform data profiling, cleansing, transformation, and validation activities.
  • Optimize Spark jobs for performance, scalability, and resource utilization.
  • Work closely with Data Scientists to prepare datasets for analytics and ML use cases.
  • Collaborate with Product Owners, Business Analysts, Architects, and cross-functional teams.
  • Monitor, troubleshoot, and resolve production data pipeline issues.
  • Participate in code reviews and implement engineering best practices.
  • Create technical documentation and maintain data engineering standards.
  • Support continuous improvement of enterprise data platforms and engineering processes.
  • Participate in Agile/Scrum ceremonies including sprint planning, backlog grooming, stand-ups, and retrospectives.

Skills

PySpark
Python
Cloudera CDP
Distributed data processing
ETL/ELT pipelines
Big Data
Hadoop ecosystem
SQL
Data modelling
Data quality
Data governance
Git
CI/CD for data
REST APIs
Structured data
Cloud-native data platforms
Machine learning data pipelines
Docker
Kubernetes
DevOps for Data Eng

Tools

CDP
Hadoop
Spark
Airflow
Docker
Kubernetes

Job description

We are looking for an experienced Senior Data Engineer with strong expertise in PySpark, Python, and Cloudera Data Platform (CDP) to join a high-performing Data Engineering team supporting enterprise-scale Digital Products & Transaction Banking initiatives.

The ideal candidate will have extensive experience designing, building, and optimizing large-scale data pipelines within modern Big Data ecosystems.

This role requires strong technical expertise in distributed data processing, cloud-native data platforms, data quality, and enterprise data engineering best practices.

The successful candidate will work closely with Data Architects, Data Scientists, Analytics teams, Product Owners, and Business Stakeholders to deliver scalable, secure, and high-performance data solutions that power business intelligence, analytics, and machine learning initiatives.

Key Responsibilities
  • Design, develop, and maintain scalable, high-performance data pipelines using PySpark and Python .
  • Build robust batch and distributed data processing solutions using Cloudera Data Platform (CDP) .
  • Develop and optimize ETL/ELT pipelines for structured and unstructured enterprise datasets.
  • Design scalable data ingestion frameworks from multiple enterprise data sources.
  • Ensure high data quality, integrity, governance, and availability across enterprise platforms.
  • Perform data profiling, cleansing, transformation, and validation activities.
  • Optimize Spark jobs for performance, scalability, and resource utilization.
  • Work closely with Data Scientists to prepare datasets for analytics and machine learning use cases.
  • Collaborate with Product Owners, Business Analysts, Architects, and cross-functional engineering teams.
  • Monitor, troubleshoot, and resolve production data pipeline issues.
  • Participate in code reviews and implement engineering best practices.
  • Create technical documentation and maintain data engineering standards.
  • Support continuous improvement of enterprise data platforms and engineering processes.
  • Participate in Agile/Scrum ceremonies including sprint planning, backlog grooming, stand-ups, and retrospectives.
Required Technical Skills
  • 8–10+ years of experience in Data Engineering.
  • Strong hands-on expertise in Python .
  • Strong hands-on expertise in PySpark .
  • Extensive experience with Cloudera Data Platform (CDP) .
  • Strong understanding of distributed data processing frameworks.
  • Experience building enterprise-scale ETL/ELT pipelines.
  • Strong knowledge of Big Data technologies.
  • Experience with Hadoop ecosystem technologies.
  • Strong SQL programming and query optimization skills.
  • Experience with data modelling and data transformation techniques.
  • Knowledge of data quality, validation, and governance principles.
  • Experience working with Git and version control systems.
  • Understanding of CI/CD practices for data engineering.
  • Strong understanding of REST APIs and data integration patterns.
  • Experience working with structured, semi-structured, and unstructured datasets.
  • Nice to Have Experience working with cloud-native data platforms.
  • Exposure to machine learning data pipelines.l>
  • Knowledge of feature engineering and data preparation for AI/ML workloads.
  • Experience with workflow orchestration tools.
  • Exposure to containerization technologies such as Docker and Kubernetes.
  • Experience with DevOps practices for Data Engineering.
Required Competencies
  • Strong analytical and problem-solving skills.
  • Excellent communication and stakeholder management skills.
  • Ability to work in fast-paced Agile delivery environments.
  • Strong ownership mindset with focus on quality and delivery.
  • Ability to collaborate effectively with business and technical stakeholders.
  • Strong debugging and performance optimization capabilities.
  • Ability to manage multiple priorities and deliver within tight timelines.
Preferred Domain Experience

Banking Financial Services Digital Products Transaction Banking Enterprise Data Platforms

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer - PySpark, Python & Cloudera (CDP)
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSSTech Group • Dubai

On-site
AED 360,000 - 520,000
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSS Group • Dubai

On-site
AED 350,000 - 520,000
Senior Data Engineer - PySpark, Python & CDP Expert
Senior Data Engineer - PySpark, Python & CDP Expert

GSSTech Group • Dubai

On-site
AED 360,000 - 520,000
Data Engineer - ETL/PySpark (Banking Domain)
Data Engineer - ETL/PySpark (Banking Domain)

GSS Group • Dubai

On-site
AED 300,000 - 520,000
Data Engineer (PySpark)
Data Engineer (PySpark)

Black Pearl Consult • Abu Dhabi

On-site
AED 260,000 - 380,000
Senior Data Engineer - PySpark, Python & CDP Expert
Senior Data Engineer - PySpark, Python & CDP Expert

GSSTech Group • United Arab Emirates

On-site
AED 250,000 - 360,000
Senior Data Platform Engineer
Senior Data Platform Engineer

BlackStone eIT • Dubai

On-site
Paid Time Off
Performance Bonus
Training & Development
Senior Data Platform Engineer
Senior Data Platform Engineer

BlackStone eIT • Dubai

On-site
AED 250,000 - 350,000
Paid Time Off
Performance Bonus
Training & Development
Data Engineer
Data Engineer

Wasl Group • Dubai

On-site
AED 150,000 - 200,000
Senior PySpark Data Engineer - ETL & Data Warehousing
Senior PySpark Data Engineer - ETL & Data Warehousing

GSS Group • Dubai

On-site
AED 300,000 - 520,000