Data Engineer

cloudandthings.io

Wes-Kaap

On-site

ZAR 720,000 - 1,100,000

Full time

46 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Flexible and supportive work culture
Career advancement opportunities
Ongoing learning and development

Job summary

cloudandthings.io is seeking a Data Engineer to design, build, and operate modern, cloud-native data platforms across AWS and Azure, leveraging Databricks, Microsoft Fabric, and AWS services. You will work across ingestion, transformation, modeling and serving to enable real-time analytics and AI use cases.

We welcome engineers from junior to principal levels and value ownership, rigorous coding practices, and the ability to translate complex requirements into production-grade data solutions

Qualifications

  • Bachelor's degree in Engineering, Computer Science, or related field.
  • Proven track record designing and implementing data solutions.
  • Hands-on with Azure and/or AWS cloud infrastructure and services.
  • Experience with at least one lakehouse platform (Databricks or Fabric).
  • Knowledge of Docker/Containers/Kubernetes/CICD pipelines for data.

Responsibilities

  • Design, build, and operate cloud-native data platforms.
  • Cover full data lifecycle: ingestion, transformation, modelling and serving for real-time analytics and AI use cases.
  • Develop scalable ETL/ELT pipelines using Databricks and Fabric.
  • Implement data security, governance, and compliance across platforms.
  • Collaborate with clients and teams to deliver production-grade data solutions.

Skills

Python
SQL
Databricks
Microsoft Fabric
AWS
Azure
Data modelling
ETL/ELT
CI/CD
Terraform
Linux
Git

Education

Bachelor's degree in Engineering, Computer Science, or related field

Tools

Databricks
Fabric
Terraform
Git
CI/CD pipelines

Job description

At cloudandthings.io, we are an engineering-led consultancy focused on building modern Data and AI Platforms for enterprise clients.

We don't just design architectures; we build real systems that power analytics, decision-making, and AI at scale. Our team acts as a force multiplier, accelerating delivery, improving quality, and unlocking measurable business value. We operate with high standards, strong ownership, and a bias for execution.

Who thrives here

We work with smart people who think from first principles, challenge conventional thinking, and take deep ownership and accountability for what they deliver. We value engineers who don't simply identify problems, but take responsibility for solving them and seeing things through.

If you care about solving meaningful problems and building production-grade systems, you'll fit right in.

Overview

As a Data Engineer, you will design, build, and operate modern, cloud-native data platforms across AWS and Azure, leveraging Databricks, Microsoft Fabric, and AWS native services. You will work across the full data lifecycle: ingestion, transformation, modelling, and serving, enabling real-time analytics, reporting, and AI use cases.

We’re looking for Data Engineers across all levels (Junior to Principal) to join our growing Data, Analytics and AI capability.

Key Responsibilities

While the list below is long, an ideal candidate should have working knowledge and experience covering many of the tools and services. The requirements for each project differ over time, and these skills provide an overview of what may typically be required of a Data Engineer.

Software Engineering Foundations
  • Strong grounding in software engineering fundamentals (data structures, algorithms, design patterns).
  • Proficiency in Python and SQL (additional languages advantageous).
  • Experience with Git, CI/CD pipelines, and modern development practices.
  • Familiarity with Terraform or Bicep for Infrastructure as Code, including the Databricks Terraform provider and Fabric deployment automation.
  • Comfortable working in Linux-based environments.
  • Packaging and promoting data assets as code: Databricks Asset Bundles, Databricks Repos, Fabric Git integration and deployment pipelines.
Data Ingestion and Streaming
  • Build scalable ingestion pipelines across hybrid and cloud environments.
  • Real-time streaming: AWS Kinesis / MSK (Kafka), Azure Event Hubs / Kafka, Databricks Structured Streaming, Fabric Eventstream and Real-Time Intelligence (Eventhouse / KQL databases, Activator).
  • Incremental and file-based ingestion: Databricks Auto Loader, Lakeflow Connect, Fabric Dataflows Gen2 and Copy Job.
  • Batch ingestion: AWS DataSync, DMS, Azure Data Factory / Synapse Pipelines / Fabric Data Pipelines.
  • Integration via APIs, JDBC/ODBC, and CDC pipelines, including Fabric Mirroring (Azure SQL, Cosmos DB, Snowflake, Databricks Unity Catalog) and Delta Sharing.
Storage, Lakehouse and Fabric
  • Design and manage data lakes using: Amazon S3, Azure Data Lake Storage Gen2 (ADLS).
  • Implement lakehouse architectures using:
  • Databricks — Delta Lake, Unity Catalog (catalogs, schemas, external locations, storage credentials, volumes), managed vs external tables, streaming tables and materialised views, Lakehouse Federation, Delta Sharing.
  • Microsoft Fabric — OneLake (including shortcuts to S3/ADLS and other Fabric items), Lakehouse, Warehouse, Fabric SQL Database, workspace and domain design, capacity (F SKU) sizing and CU management.
  • Table optimisation and maintenance: partitioning, liquid clustering, Z-ordering, OPTIMIZE / VACUUM, V-Order and statistics in Fabric, small-file management, time travel and retention policies.
  • Work with modern data formats and open standards: Delta Lake, Parquet, Avro, JSON, Apache Iceberg / Delta UniForm for cross-engine interoperability.
  • Experience with: Relational databases (Postgres, SQL Server, Aurora), NoSQL (DynamoDB, Cosmos DB), Caching (Redis).
Data Processing and Transformation
  • Build scalable ETL/ELT pipelines using:
  • Databricks — PySpark and Spark SQL, Delta Live Tables / Lakeflow Declarative Pipelines, Databricks Workflows, serverless and job compute, Photon, cluster policies, dbt on Databricks.
  • Microsoft Fabric — Fabric Data Engineering (Spark notebooks, Spark job definitions, environments and library management), Fabric Warehouse T-SQL, stored procedures, and Dataflows Gen2.
  • Implement medallion architecture (Bronze/Silver/Gold) and dimensional models (star schemas, SCD handling) on the lakehouse.
  • Develop reusable, testable, and production-grade data pipelines: modular notebooks and Python packages, parameterisation, unit and integration testing, idempotent and restartable loads.
  • Performance and cost tuning of Spark workloads: skew and shuffle management, caching strategies, right-sizing compute, autoscaling.
Analytics and AI Enablement
  • Design platforms that support Business Intelligence, advanced analytics, machine learning and AI use cases.
  • Work with:
  • Databricks — Databricks SQL warehouses, AI/BI dashboards and Genie, MLflow, Unity Catalog model registry, Model Serving, Mosaic AI (Vector Search, agent and RAG patterns), feature engineering in Unity Catalog.
  • Microsoft Fabric — semantic models and Direct Lake mode, Power BI report and dataset design, DAX fundamentals, Fabric Data Science (notebooks, MLflow experiments), Fabric Data Agents and Copilot experiences.
  • AWS — Amazon Redshift / Athena, SageMaker and Bedrock where relevant.
  • Support feature engineering, data science workflows, and real-time decisioning systems.
  • Implement data quality, observability, and lineage frameworks: DLT expectations, Unity Catalog lineage and system tables, Fabric monitoring hub and lineage view, alerting on freshness and volume anomalies.
Security, Governance & Compliance
  • Implement secure, enterprise-grade data platforms: AWS IAM / Azure Entra ID (AAD), RBAC, Managed Identities, service principals.
  • Governance:
  • Databricks Unity Catalog — three-level namespace design, grants and privilege inheritance, row filters and column masks, attribute-based access, audit logs and system tables.
  • Microsoft Fabric — workspace roles and item permissions, OneLake security, sensitivity labels, domains, and integration with Microsoft Purview.
  • AWS Lake Formation — fine-grained access control over S3-based lakes.
  • Networking: VPC / VNets, Private Endpoints, Direct Connect / ExpressRoute, Databricks secure cluster connectivity / Private Link, Fabric private links and managed VNets.
  • Encryption: KMS / Key Vault / TLS, customer-managed keys, secret scopes.
Orchestration and Operations
  • Build orchestrated pipelines using: Databricks Workflows (job dependencies, retries, notifications), AWS Step Functions / MWAA (Airflow), ADF / Synapse / Fabric Data Pipelines and notebook scheduling.
  • Monitoring & observability: cloud-native monitoring tools (CloudWatch, Azure Monitor, Log Analytics), Databricks system tables and query history, Fabric Capacity Metrics app and monitoring hub.
  • Apply best practices across reliability, performance optimisation, and cost optimisation (FinOps): Databricks compute policies, serverless vs classic trade-offs, tagging and chargeback, Fabric capacity smoothing, bursting, and pause/resume strategies.
Requirements
  • Bachelor's degree in Engineering, Computer Science, or related field.
  • Proven track record of designing and implementing data solutions.
  • Knowledge of and experience with Azure and/or AWS Cloud infrastructure and services.
  • Hands‑on delivery experience with at least one lakehouse platform (Databricks or Microsoft Fabric), and an appetite to build depth in the other.
  • Certifications, such as:
  • Microsoft — Fabric Data Engineer Associate (DP-700), Fabric Analytics Engineer Associate (DP-600), Azure Data Engineer Associate (DP-203), Azure Solutions Architect Expert (AZ-305).
  • AWS — Data Engineer Associate, Solutions Architect Associate / Professional.
  • Any other data-related experience, e.g. working with Hadoop, databases, analytics software, etc.
  • Experience with Docker/Containers/Kubernetes/CICD pipelines for data.
  • Knowledge of data security and compliance standards.
  • Willingness to learn and expand knowledge related to Cloud and Data Technologies.
  • Strong problem-solving and analytical skills.
  • Self-organising with the ability to prioritise and manage multiple tasks simultaneously.
  • Excellent verbal and written communication skills.
  • Ability to work collaboratively with clients and team members.
  • Willingness to travel to clients as and when required.
What We Offer
  • A culture of engineering and an environment where ideas are heard and builders can build.
  • Competitive compensation and bonus structure.
  • A flexible and supportive work environment that values diversity, work-life balance, and personal growth.
  • Opportunities for career advancement and ongoing professional development.
  • Ongoing learning and development opportunities to enhance your skills.
  • Engaging with cutting-edge technologies and awesome client projects.
  • Access to a talented team of professionals and mentors.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

cloudandthings.io • City of Johannesburg Metropolitan Municipality

On-site
ZAR 600,000 - 1,200,000
Competitive compensation
Flexible work environment
Career development
+1
Data Engineer (Analytics & Data Platform)
Data Engineer (Analytics & Data Platform)

ATS Client • Cape Town

Hybrid
ZAR 1,200,000 - 1,800,000
Training budget
Flexible working arrangements
Career development
Data Engineer
Data Engineer

ATS Client • Johannesburg

On-site
ZAR 850,000 - 1,250,000
Data Engineering Lead
Data Engineering Lead

Blue Pearl HQ • Johannesburg

On-site
ZAR 1,100,000 - 1,900,000
Senior Data Engineer
Senior Data Engineer

BETSoftware • Johannesburg

On-site
ZAR 800,000 - 1,200,000
Cloud and Data Engineer
Cloud and Data Engineer

Jobtailor • KwaZulu-Natal

Hybrid
ZAR 700,000 - 1,200,000
Data Engineer
Data Engineer

Cloud Direct • Wes-Kaap

On-site
ZAR 600,000 - 900,000
Uncapped annual leave
Group Life Cover
Fitness Cash Contribution
+5
Senior Data Fabric Data Engineer
Senior Data Fabric Data Engineer

Belay Talent Solutions • Midrand

Hybrid
ZAR 900,000 - 1,350,000
Senior Data Engineer (Azure / Microsoft Fabric) - Banking Sector
Senior Data Engineer (Azure / Microsoft Fabric) - Banking Sector

ATS Client • Johannesburg

On-site
ZAR 900,000 - 1,500,000
Cloud Engineer (Contract)
Cloud Engineer (Contract)

The Focus Group • Sandton

On-site
ZAR 800,000 - 1,200,000