Senior Data Engineer – Big Data & Cloudera

Sabenza IT & Recruitment

Johannesburg

On-site

ZAR 600,000 - 1,200,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sabenza IT & Recruitment seeks a data engineer to join a high-performing data engineering environment in Johannesburg, South Africa. You will design, develop and optimize enterprise-scale data solutions, focusing on CDP and Hadoop ecosystems, Spark, PySpark, NiFi, Sqoop, and SQL-driven analytics.

You will build robust ETL pipelines, process large datasets, and support analytics initiatives with Python and shell scripting in a Linux/Unix environment.

Qualifications

  • 7–8 years of solid hands-on experience as a platform and data engineer (intermediate to senior level).
  • Design, develop and maintain scalable Big Data and ETL data pipelines.
  • Work extensively with the Cloudera Data Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.

Responsibilities

  • Design, develop and maintainscalable Big Data and ETL data pipelines.
  • Work extensively with CDP and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts, Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuous improvement of data engineering standards, processes and platforms.

Skills

Big Data pipelines
ETL development
Data integration
SQL queries
Python scripting
Shell scripting
Unix/Linux
Data quality
Collaboration

Tools

Cloudera CDP
Hadoop ecosystem
Apache Spark
PySpark
Apache NiFi
Sqoop
HDFS
Hive
Impala

Job description

Johannesburg, South Africa

We are looking for a skilled Data Engineer tojoin a high-performing data engineering environment and contribute to thedesign, development, integration and optimization of enterprise-scale datasolutions.

The ideal candidate will have strong hands-onexperience across the Cloudera Data Platform (CDP) and the broader Hadoopecosystem , with proven expertise in building robust ETL pipelines,processing large datasets and supporting data analytics initiatives.

This is an exciting opportunity for a dataengineering professional who enjoys working with Big Data technologies ,solving complex data challenges and building scalable solutions that enablesmarter business decisions.

Key Responsibilities

  • Design, develop and maintainscalable Big Data and ETL data pipelines .
  • Work extensively with the ClouderaData Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize dataprocessing solutions using Apache Spark and PySpark .
  • Build and manage data ingestionpipelines using Apache NiFi and Sqoop .
  • Work with HDFS, Hive andImpala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineeringsolutions using Python and Shell scripting.
  • Integrate and process data fromenterprise data sources, including Oracle .
  • Develop, maintain and optimizeETL processes to support business and analytical requirements.
  • Monitor data pipelines andscheduled workloads using Control-M .
  • Perform troubleshooting,performance tuning and root-cause analysis across data processingenvironments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineeringprocesses.
  • Support data quality, dataintegrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts,Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuousimprovement of data engineering standards, processes and platforms.
Requirements

7–8 years of solid hands-on experience as a platform and data engineer (intermediate to senior level).

  • Design, develop and maintainscalable Big Data and ETL data pipelines .
  • Work extensively with the ClouderaData Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize dataprocessing solutions using Apache Spark and PySpark .
  • Build and manage data ingestionpipelines using Apache NiFi and Sqoop .
  • Work with HDFS, Hive andImpala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineeringsolutions using Python and Shell scripting.
  • Integrate and process data fromenterprise data sources, including Oracle .
  • Develop, maintain and optimizeETL processes to support business and analytical requirements.
  • Monitor data pipelines andscheduled workloads using Control-M .
  • Perform troubleshooting,performance tuning and root-cause analysis across data processingenvironments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineeringprocesses.
  • Support data quality, dataintegrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts,Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuousimprovement of data engineering standards, processes and platforms.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer – Cloudera CDP & Big Data Pipelines
Senior Data Engineer – Cloudera CDP & Big Data Pipelines

Sabenza IT & Recruitment • Johannesburg

On-site
ZAR 600,000 - 1,200,000
Data Engineering Lead
Data Engineering Lead

Blue Pearl PTY • Johannesburg

On-site
ZAR 1,200,000 - 1,900,000
Senior Data Engineer - JHB
Senior Data Engineer - JHB

Hire Resolve • Johannesburg

On-site
ZAR 800,000 - 1,100,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

Sabenza IT & Recruitment • Cape Town

On-site
ZAR 900,000 - 1,700,000
Senior Data Solutions Engineer
Senior Data Solutions Engineer

Future Fit • Johannesburg

Hybrid
ZAR 900,000 - 1,500,000
Competitive compensation package
Twice-yearly salary increases
Employee wellness programs
+1
Data Engineer
Data Engineer

PBT Group • Cape Town

On-site
ZAR 900,000 - 1,300,000
Data Engineer
Data Engineer

CVQuest • Centurion

Hybrid
ZAR 600,000 - 900,000
Data Engineer
Data Engineer

Hire Resolve • Sandton

On-site
ZAR 700,000 - 1,000,000
Strong Intermediate Data Engineer
Strong Intermediate Data Engineer

Blue Pearl PTY • Johannesburg

Hybrid
ZAR 420,000 - 660,000
2x SENIOR DATA ENGINEER – JOHANNESBURG – GAUTENG
2x SENIOR DATA ENGINEER – JOHANNESBURG – GAUTENG

Tych Business Solutions • Johannesburg

On-site
ZAR 600,000 - 900,000