Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSSTech Group

Dubai

On-site

AED 360,000 - 520,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GSSTech Group is seeking a Senior Data Engineer to design and optimize large-scale data pipelines using PySpark, Python, and Cloudera Data Platform (CDP). You will collaborate with data architects, scientists, and stakeholders to deliver scalable data solutions for BI, analytics, and ML initiatives.

The ideal candidate has 8-10+ years in data engineering, strong SQL, data modeling, and experience with CDP and Hadoop ecosystems.

Qualifications

  • 8-10+ years of Data Engineering experience.
  • Proficient in Python and PySpark.
  • Extensive CDP experience and big data orchestration.
  • Strong SQL and data modeling abilities.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using PySpark and Python.
  • Build and optimize batch and distributed data processing with CDP.
  • Develop and optimize ETL/ELT pipelines for structured and unstructured data.
  • Design scalable data ingestion frameworks from multiple enterprise sources.
  • Ensure data quality, governance, and availability across platforms.
  • Perform profiling, cleansing, transformation, and validation of data.
  • Optimize Spark jobs for performance and resource use.
  • Collaborate with Data Scientists on analytics-ready datasets.
  • Coordinate with Product Owners and engineers in Agile ceremonies.
  • Troubleshoot production data pipelines and resolve issues.
  • Contribute to code reviews and engineering best practices.
  • Create technical docs and data engineering standards.
  • Support continuous improvement of data platforms and processes.

Skills

PySpark
Python
SQL
Data modeling
Data quality
Distributed processing
Agile/Scrum

Tools

CDP (Cloudera Data Platform)
Hadoop
Git
Kubernetes

Job description

We are looking for an experienced Senior Data Engineer with strong expertise in PySpark, Python, and Cloudera Data Platform (CDP) to join a high-performing Data Engineering team supporting enterprise-scale Digital Products & Transaction Banking initiatives.

The ideal candidate will have extensive experience designing, building, and optimizing large-scale data pipelines within modern Big Data ecosystems. This role requires strong technical expertise in distributed data processing, cloud-native data platforms, data quality, and enterprise data engineering best practices.

The successful candidate will work closely with Data Architects, Data Scientists, Analytics teams, Product Owners, and Business Stakeholders to deliver scalable, secure, and high-performance data solutions that power business intelligence, analytics, and machine learning initiatives.

Key Responsibilities
  • Design, develop, and maintain scalable, high-performance data pipelines using PySpark and Python.
  • Build robust batch and distributed data processing solutions using Cloudera Data Platform (CDP).
  • Develop and optimize ETL/ELT pipelines for structured and unstructured enterprise datasets.
  • Design scalable data ingestion frameworks from multiple enterprise data sources.
  • Ensure high data quality, integrity, governance, and availability across enterprise platforms.
  • Perform data profiling, cleansing, transformation, and validation activities.
  • Optimize Spark jobs for performance, scalability, and resource utilization.
  • Work closely with Data Scientists to prepare datasets for analytics and machine learning use cases.
  • Collaborate with Product Owners, Business Analysts, Architects, and cross-functional engineering teams.
  • Monitor, troubleshoot, and resolve production data pipeline issues.
  • Participate in code reviews and implement engineering best practices.
  • Create technical documentation and maintain data engineering standards.
  • Support continuous improvement of enterprise data platforms and engineering processes.
  • Participate in Agile/Scrum ceremonies including sprint planning, backlog grooming, stand-ups, and retrospectives.
Required Technical Skills
  • 8-10+ years of experience in Data Engineering.
  • Strong hands-on expertise in Python.
  • Strong hands-on expertise in PySpark.
  • Extensive experience with Cloudera Data Platform (CDP).
  • Strong understanding of distributed data processing frameworks.
  • Experience building enterprise-scale ETL/ELT pipelines.
  • Strong knowledge of Big Data technologies.
  • Experience with Hadoop ecosystem technologies.
  • Strong SQL programming and query optimization skills.
  • Experience with data modelling and data transformation techniques.
  • Knowledge of data quality, validation, and governance principles.
  • Experience working with Git and version control systems.
  • Understanding of CI/CD practices for data engineering.
  • Strong understanding of REST APIs and data integration patterns.
  • Experience working with structured, semi-structured, and unstructured datasets.
Nice to Have
  • Experience working with cloud-native data platforms.
  • Exposure to machine learning data pipelines.
  • Knowledge of feature engineering and data preparation for AI/ML workloads.
  • Experience with workflow orchestration tools.
  • Exposure to containerization technologies such as Docker and Kubernetes.
  • Experience with DevOps practices for Data Engineering.
Required Competencies
  • Strong analytical and problem-solving skills.
  • Excellent communication and stakeholder management skills.
  • Ability to work in fast-paced Agile delivery environments.
  • Strong ownership mindset with focus on quality and delivery.
  • Ability to collaborate effectively with business and technical stakeholders.
  • Strong debugging and performance optimization capabilities.
  • Ability to manage multiple priorities and deliver within tight timelines.
Preferred Domain Experience
  • Banking
  • Financial Services
  • Digital Products
  • Transaction Banking
  • Enterprise Data Platforms
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer - PySpark, Python & Cloudera (CDP)
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSS Group • Dubai

On-site
AED 350,000 - 520,000
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)
Sr. Data Engineer - PySpark, Python & Cloudera (CDP)

GSSTech Group • United Arab Emirates

On-site
AED 250,000 - 360,000
Senior Data Engineer - PySpark, Python & CDP Expert
Senior Data Engineer - PySpark, Python & CDP Expert

GSSTech Group • Dubai

On-site
AED 360,000 - 520,000
Data Engineer (PySpark)
Data Engineer (PySpark)

Black Pearl Consult • Abu Dhabi

On-site
AED 260,000 - 380,000
Data Engineer - Databricks
Data Engineer - Databricks

Flentas • Dubai

On-site
AED 420,000 - 660,000
Senior Data Engineer - PySpark, Python & CDP Expert
Senior Data Engineer - PySpark, Python & CDP Expert

GSSTech Group • United Arab Emirates

On-site
AED 250,000 - 360,000
Data Engineer
Data Engineer

Wasl Group • Dubai

On-site
AED 150,000 - 200,000
Senior Data Platform Engineer
Senior Data Platform Engineer

BlackStone eIT • Dubai

On-site
AED 250,000 - 350,000
Paid Time Off
Performance Bonus
Training & Development
Senior Data Platform Engineer
Senior Data Platform Engineer

BlackStone eIT • Dubai

On-site
Paid Time Off
Performance Bonus
Training & Development
Senior Data Engineer
Senior Data Engineer

Innovo Group • Dubai

On-site
AED 240,000 - 360,000