Junior Data Engineer

Cls

Pretoria

On-site

ZAR 420,000 - 540,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Own transport
Driver's license

Job summary

Cls in Pretoria is seeking a Junior Data Engineer to manage an established Microsoft Fabric data platform and develop new data engineering capabilities. You will own end-to-end data pipelines, from ingestion to BI-ready datasets, across batch and near-real-time processing.

The role demands hands-on expertise with Fabric components and strong problem-solving across data layers. The successful candidate will collaborate with analytics teams to deliver robust data models and scalable solutions,

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Data Engineering or a related field.
  • Relevant Microsoft Fabric and/or Azure data certification.
  • 0-1 years of practical data engineering experience, including strong recent hands-on experience with Microsoft Fabric.
  • Strong hands-on experience with Microsoft Fabric, including OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
  • Strong SQL Server and T-SQL capability, including complex query development, schema design, indexing and performance optimisation.
  • Practical experience developing, maintaining and supporting production ETL/ELT pipelines.
  • Experience integrating and extracting data from REST/SOAP APIs, databases, flat files and other structured or unstructured data sources.
  • Proficiency in data transformation using SQL and Python/PySpark, with an understanding of scalable data processing practices.
  • Practical experience in data warehousing, dimensional modelling, incremental loading, orchestration and schema evolution.
  • Experience with source control, CI/CD and deployment practices using Git, Azure DevOps or equivalent tools.
  • Experience supporting Power BI and other downstream analytical or reporting requirements.

Responsibilities

  • Take ownership of the existing Microsoft Fabric data engineering environment and become productive quickly following handover.
  • Design, develop, maintain and orchestrate reliable batch and near-real-time data ingestion pipelines using appropriate Microsoft Fabric capabilities.
  • Extract and ingest data from structured and unstructured sources, including REST APIs, SOAP APIs, databases and flat files.
  • Develop robust data transformation logic using SQL, Python/PySpark, Fabric notebooks and Dataflows Gen2, as appropriate.
  • Implement incremental loading, retry mechanisms, logging, monitoring and alerting to support data integrity and pipeline reliability.
  • Troubleshoot and resolve pipeline failures and data processing issues efficiently.
  • Optimise data pipelines and processing workloads for performance, scalability and cost-effectiveness.

Skills

Microsoft Fabric
SQL Server
T-SQL
Python
PySpark
REST APIs
SOAP APIs
Data Warehousing
ETL/ELT
Power BI integration

Education

Bachelor's degree in Computer Science / Information Systems / Data Engineering

Tools

Git
Azure DevOps
CI/CD
OneLake
Lakehouse
Data Factory
Power BI

Job description

Introduction:
We are seeking an experienced Junior Data Engineer to join a company based in Pretoria and manage and develop an established Microsoft Fabric data platform. The successful candidate will ensure reliable day-to-day operations while delivering new data engineering requirements. This is a hands-on role requiring strong technical capability, independent problem-solving and the ability to work across data ingestion, transformation, modelling, warehousing and business intelligence integration.
Job Purpose:
To take ownership of an established Microsoft Fabric data platform, ensuring reliable operations and delivery of new data engineering requirements.
REQUIREMENTS
Minimum Education (essential)

  • Bachelor's degree in Computer Science, Information Systems, Data Engineering or a related field.
  • Relevant Microsoft Fabric and/or Azure data certification.

Minimum applicable experience (years):

  • 0-1 years of practical data engineering experience, including strong recent hands-on experience with Microsoft Fabric.

Required nature of experience:

  • Strong hands-on experience with Microsoft Fabric, including OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
  • Strong SQL Server and T-SQL capability, including complex query development, schema design, indexing and performance optimisation.
  • Practical experience developing, maintaining and supporting production ETL/ELT pipelines.
  • Experience integrating and extracting data from REST/SOAP APIs, databases, flat files and other structured or unstructured data sources.
  • Proficiency in data transformation using SQL and Python/PySpark, with an understanding of scalable data processing practices.
  • Practical experience in data warehousing, dimensional modelling, incremental loading, orchestration and schema evolution.
  • Experience with troubleshooting pipeline failures, data quality issues and performance bottlenecks, including the ability to restore service efficiently.
  • Experience with source control, CI/CD and deployment practices using Git, Azure DevOps or equivalent tools.
  • Experience supporting Power BI and other downstream analytical or reporting requirements.
  • Demonstrated ability to take ownership of an existing technical environment with limited hand-holding.
  • Strong documentation, communication and stakeholder engagement skills.

Skills and Knowledge (essential):

  • Microsoft Fabric: OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
  • SQL Server / T-SQL
  • Python / PySpark
  • REST/SOAP APIs and structured/unstructured data ingestion.
  • ETL/ELT, incremental loading, orchestration and scheduling.
  • Dimensional modelling, medallion architecture, schema evolution and data warehousing.
  • Power BI integration and understanding of downstream analytical requirements.
  • Git / Azure DevOps, CI/CD and environment deployment practices.
  • Monitoring, data quality, performance optimisation, security and operational support.

Other:

  • Proficient in Afrikaans and English.
  • Own transport and valid driver’s license.


KEY PERFORMANCE AREAS AND OBJECTIVES
Fabric Data Engineering and Pipeline Development

  • Take ownership of the existing Microsoft Fabric data engineering environment and become productive quickly following handover.
  • Design, develop, maintain and orchestrate reliable batch and near-real-time data ingestion pipelines using appropriate Microsoft Fabric capabilities.
  • Extract and ingest data from structured and unstructured sources, including REST APIs, SOAP APIs, databases and flat files.
  • Develop robust data transformation logic using SQL, Python/PySpark, Fabric notebooks and Dataflows Gen2, as appropriate.
  • Implement incremental loading, retry mechanisms, logging, monitoring and alerting to support data integrity and pipeline reliability.
  • Troubleshoot and resolve pipeline failures and data processing issues efficiently.
  • Optimise data pipelines and processing workloads for performance, scalability and cost-effectiveness.

Data Architecture and Platform Management

  • Design, manage and evolve scalable data architectures using Microsoft Fabric, OneLake, Lakehouse, Warehouse and SQL Server.
  • Maintain appropriate data-layering and medallion architecture principles, where applicable, with clear movement from raw to curated data.
  • Develop and maintain robust schema designs, indexes, partitioning and query strategies to support analytical and operational workloads.
  • Manage schema evolution and version control to maintain consistency and minimise disruption to downstream consumers.
  • Maintain metadata, data dictionaries, architecture documentation and technical documentation to improve supportability and reduce key-person dependency.
  • Define and maintain appropriate role-based access and security controls.

Data Warehousing and BI Integration

  • Build and maintain analytical data stores using Microsoft Fabric Warehouse and/or Lakehouse patterns.
  • Apply appropriate data-loading, partitioning, storage optimisation and query-performance practices.
  • Develop and maintain stable, well-modelled datasets for Power BI and other analytical consumers.
  • Work with reporting and analytical teams to investigate and resolve data-related issues.
  • Ensure data structures and outputs support downstream reporting and business intelligence requirements.

Data Modelling and Standards

  • Develop and maintain conceptual, logical and physical data models.
  • Apply dimensional modelling techniques, including star and snowflake schemas, to support analytics and reporting.
  • Apply appropriate normalisation and relational modelling techniques for operational and analytical workloads.
  • Ensure consistency of data models across systems.
  • Manage schema versioning and evolution without unnecessarily disrupting downstream consumers.
  • Apply agreed data engineering standards and modelling principles consistently.

Ownership, Reporting and Communication

  • Work independently and take end-to-end ownership of assigned data engineering deliverables, incidents and production issues.
  • Provide clear and timely updates regarding progress, risks, dependencies and blockers.
  • Engage directly with technical and business stakeholders to clarify requirements and agree practical solutions.
  • Explain technical concepts and trade-offs in a manner appropriate to the relevant stakeholder.
  • Maintain practical technical documentation, including runbooks, architecture notes, change logs and release notes.
  • Take accountability for the successful delivery and operational support of assigned solutions.

Automation, Monitoring and Optimisation

  • Automate recurring data engineering and operational activities where practical.
  • Implement monitoring and alerting to identify data quality issues, pipeline failures and abnormal processing behaviour.
  • Analyse and optimise query, notebook and pipeline performance across SQL Server and Microsoft Fabric.
  • Monitor capacity and resource utilisation and contribute to scalability and cost-control decisions.
  • Deploy solutions using appropriate CI/CD and controlled deployment practices.

Security and Best Practices

  • Apply data security best practices, including secure authentication, least-privilege access and appropriate encryption.
  • Ensure data engineering solutions comply with applicable data governance policies and regulatory requirements.
  • Apply sound engineering practices relating to recoverability, auditability, supportability and controlled change.
  • Protect confidential and sensitive business information.

Contribution to the Team

  • Collaborate with developers, data analysts, data scientists and business stakeholders to understand requirements and deliver practical solutions.
  • Support effective handover and knowledge transfer to reduce key-person dependency within the data environment.
  • Share technical knowledge and contribute to continuous improvement of team practices and the data environment.
  • Provide guidance and support to junior team members where required.
  • Remain accountable for the quality, reliability and timeliness of own deliverables.

Quality Management and Compliance

  • Document data processes, transformations, dependencies and architectural decisions.
  • Validate data outputs through reconciliation, data quality checks and appropriate testing before production deployment.
  • Maintain high standards of engineering quality by following agreed development, code review, testing, deployment, backup and archival practices.
  • Ensure changes are appropriately tested, documented and controlled before implementation.
  • Safeguard confidential information and data.
  • Support compliance with applicable organisational policies, standards and regulatory requirements.

Remuneration Offered
Market related

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Intermediate Data Engineer
Intermediate Data Engineer

Cls • Pretoria

On-site
ZAR 700,000 - 900,000
Junior Data Engineer
Junior Data Engineer

Creative Leadership Solutions • Pretoria

On-site
ZAR 350,000 - 550,000
Data Engineer (Analytics & Data Platform)
Data Engineer (Analytics & Data Platform)

ATS Client • Cape Town

Hybrid
ZAR 1,200,000 - 1,800,000
Training budget
Flexible working arrangements
Career development
Data Engineer (Fmcg / Food / Retail Industry Experience)
Data Engineer (Fmcg / Food / Retail Industry Experience)

Merand Recruitment • South Africa

On-site
ZAR 600,000 - 900,000
Senior Data Fabric Data Engineer
Senior Data Fabric Data Engineer

Belay Talent Solutions • Midrand

Hybrid
ZAR 900,000 - 1,350,000
Senior Data Engineer (MS SQL, Ms Fabric) - Johannesburg - up to R1.2mil per annum
Senior Data Engineer (MS SQL, Ms Fabric) - Johannesburg - up to R1.2mil per annum

ATS Client • Johannesburg

On-site
ZAR 720,000 - 1,200,000
Hybrid work model
Senior Data Engineer
Senior Data Engineer

Chosen Online Pty Ltd • Johannesburg

Hybrid
Competitive salary package (R80k - R110k per month)
Hybrid working model
Opportunity to work on challenging international projects
+2
Data and AI Engineer
Data and AI Engineer

WatersEdge Solutions • Gauteng

Hybrid
ZAR 900,000 - 1,300,000
Hybrid work model
Hands-on with Microsoft Fabric and AI
Junior Data Engineer
Junior Data Engineer

Sasso Consulting (Pty) Ltd • Cape Town

On-site
ZAR 391,000 - 469,000
Data and AI Engineer
Data and AI Engineer

WatersEdge Solutions • Johannesburg

Hybrid
ZAR 600,000 - 900,000