Senior Data Engineer – Cloudera CDP & Big Data Pipelines

Sabenza IT & Recruitment

Johannesburg

On-site

ZAR 600,000 - 1,200,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sabenza IT & Recruitment seeks a data engineer to join a high-performing data engineering environment in Johannesburg, South Africa. You will design, develop and optimize enterprise-scale data solutions, focusing on CDP and Hadoop ecosystems, Spark, PySpark, NiFi, Sqoop, and SQL-driven analytics.

You will build robust ETL pipelines, process large datasets, and support analytics initiatives with Python and shell scripting in a Linux/Unix environment.

Qualifications

  • 7–8 years of solid hands-on experience as a platform and data engineer (intermediate to senior level).
  • Design, develop and maintain scalable Big Data and ETL data pipelines.
  • Work extensively with the Cloudera Data Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.

Responsibilities

  • Design, develop and maintainscalable Big Data and ETL data pipelines.
  • Work extensively with CDP and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts, Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuous improvement of data engineering standards, processes and platforms.

Skills

Big Data pipelines
ETL development
Data integration
SQL queries
Python scripting
Shell scripting
Unix/Linux
Data quality
Collaboration

Tools

Cloudera CDP
Hadoop ecosystem
Apache Spark
PySpark
Apache NiFi
Sqoop
HDFS
Hive
Impala

Job description

Sabenza IT & Recruitment seeks a data engineer to join a high-performing data engineering environment in Johannesburg, South Africa. You will design, develop and optimize enterprise-scale data solutions, focusing on CDP and Hadoop ecosystems, Spark, PySpark, NiFi, Sqoop, and SQL-driven analytics.

You will build robust ETL pipelines, process large datasets, and support analytics initiatives with Python and shell scripting in a Linux/Unix environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer – Big Data & Cloudera
Senior Data Engineer – Big Data & Cloudera

Sabenza IT & Recruitment • Johannesburg

On-site
ZAR 600,000 - 1,200,000
Lead Data Engineer: Cloud Pipelines, Spark & AWS
Lead Data Engineer: Cloud Pipelines, Spark & AWS

Sabenza IT & Recruitment • Cape Town

On-site
ZAR 900,000 - 1,700,000
Senior Data Engineer: SAP HANA, AWS & Data Transformation
Senior Data Engineer: SAP HANA, AWS & Data Transformation

Sabenza IT & Recruitment • Pretoria

Hybrid
ZAR 1,000,000 - 1,600,000
Data Engineer: Build Scalable Cloud Pipelines
Data Engineer: Build Scalable Cloud Pipelines

Network Finance • Centurion

On-site
ZAR 600,000 - 900,000
Senior Data & BI Engineer — Azure Data, Power BI & ETL
Senior Data & BI Engineer — Azure Data, Power BI & ETL

Sabenza IT & Recruitment • Cape Town

On-site
ZAR 600,000 - 800,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

Sabenza IT & Recruitment • Cape Town

On-site
ZAR 900,000 - 1,700,000
Senior Data Engineer — Big Data & Cloud Pipelines
Senior Data Engineer — Big Data & Cloud Pipelines

Tenth Revolution Group • Gauteng

On-site
ZAR 600,000 - 900,000
Lead Data Engineer: Architect Scalable Pipelines
Lead Data Engineer: Architect Scalable Pipelines

Future Fit • Johannesburg

On-site
ZAR 900,000 - 1,300,000
Senior Cloud Data Engineer - Hybrid, Scalable Pipelines
Senior Cloud Data Engineer - Hybrid, Scalable Pipelines

E-Merge • South Africa

Hybrid
ZAR 972,000 - 1,188,000
Senior Databricks Data Engineer — Onsite in Cape Town
Senior Databricks Data Engineer — Onsite in Cape Town

Hileya - Management Consulting • Sandton

On-site
ZAR 500,000 - 700,000