Johannesburg, South Africa
We are looking for a skilled Data Engineer tojoin a high-performing data engineering environment and contribute to thedesign, development, integration and optimization of enterprise-scale datasolutions.
The ideal candidate will have strong hands-onexperience across the Cloudera Data Platform (CDP) and the broader Hadoopecosystem , with proven expertise in building robust ETL pipelines,processing large datasets and supporting data analytics initiatives.
This is an exciting opportunity for a dataengineering professional who enjoys working with Big Data technologies ,solving complex data challenges and building scalable solutions that enablesmarter business decisions.
Key Responsibilities
- Design, develop and maintainscalable Big Data and ETL data pipelines .
- Work extensively with the ClouderaData Platform (CDP) and Hadoop ecosystem.
- Develop and optimize dataprocessing solutions using Apache Spark and PySpark .
- Build and manage data ingestionpipelines using Apache NiFi and Sqoop .
- Work with HDFS, Hive andImpala for large-scale data storage, processing and querying.
- Develop complex and optimized SQL queries for data extraction, transformation and analysis.
- Develop data engineeringsolutions using Python and Shell scripting.
- Integrate and process data fromenterprise data sources, including Oracle .
- Develop, maintain and optimizeETL processes to support business and analytical requirements.
- Monitor data pipelines andscheduled workloads using Control-M .
- Perform troubleshooting,performance tuning and root-cause analysis across data processingenvironments.
- Work within Linux/Unix environments to administer, troubleshoot and automate data engineeringprocesses.
- Support data quality, dataintegrity and data availability across enterprise data platforms.
- Collaborate with Data Analysts,Developers, Architects, Business Analysts and other technology teams.
- Contribute to the continuousimprovement of data engineering standards, processes and platforms.
Requirements
7–8 years of solid hands-on experience as a platform and data engineer (intermediate to senior level).
- Design, develop and maintainscalable Big Data and ETL data pipelines .
- Work extensively with the ClouderaData Platform (CDP) and Hadoop ecosystem.
- Develop and optimize dataprocessing solutions using Apache Spark and PySpark .
- Build and manage data ingestionpipelines using Apache NiFi and Sqoop .
- Work with HDFS, Hive andImpala for large-scale data storage, processing and querying.
- Develop complex and optimized SQL queries for data extraction, transformation and analysis.
- Develop data engineeringsolutions using Python and Shell scripting.
- Integrate and process data fromenterprise data sources, including Oracle .
- Develop, maintain and optimizeETL processes to support business and analytical requirements.
- Monitor data pipelines andscheduled workloads using Control-M .
- Perform troubleshooting,performance tuning and root-cause analysis across data processingenvironments.
- Work within Linux/Unix environments to administer, troubleshoot and automate data engineeringprocesses.
- Support data quality, dataintegrity and data availability across enterprise data platforms.
- Collaborate with Data Analysts,Developers, Architects, Business Analysts and other technology teams.
- Contribute to the continuousimprovement of data engineering standards, processes and platforms.