Senior Data Engineer

cloudandthings.io

Johannesburg

On-site

ZAR 1,000,000 - 1,600,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation package
Flexible work environment
Career development and mentorship
Cutting-edge technology projects

Job summary

cloudandthings.io is seeking a Senior Data Engineer to lead the design, build and operation of cloud-native data platforms across AWS and Azure, enabling real-time analytics and AI use cases. You will own delivery across the full data lifecycle—ingestion, transformation, modelling and serving—and work hands-on while providing technical direction.

Based in Johannesburg, you will shape architecture, set engineering standards, mentor engineers, and drive production-grade solutions for enterprise

Qualifications

  • Bachelor's degree in Engineering, Computer Science, or a related field.
  • 7+ years in data or software engineering with data platforms.
  • Experience with lakehouse platforms (Databricks or Fabric).

Responsibilities

  • Lead design and delivery of cloud-native data platforms.
  • Own end-to-end solution design from discovery to production.
  • Set engineering standards and review processes.
  • Mentor engineers and contribute to hiring.
  • Collaborate with clients to shape architecture and drive outcomes.

Skills

Python
SQL
Databricks
Microsoft Fabric
Terraform
Git
CI/CD
Linux
Spark / PySpark

Education

Bachelor's degree in Engineering, Computer Science, or related field

Tools

Databricks
Microsoft Fabric
Terraform
Bicep
Git
CI/CD

Job description

About cloudandthings.io

At cloudandthings.io, we are an engineering-led consultancy focused on building modern Data and AI Platforms for enterprise clients.

We don't just design architectures; we build real systems that power analytics, decision-making, and AI at scale. We operate with high standards, strong ownership, and a bias for execution.

We are looking for people who think from first principles, challenge conventional thinking, and take deep ownership of what they deliver - engineers who don't simply identify problems, but solve them and see them through to a running system in production.

The Role

As a Senior Data Engineer, you will lead the design, build, and operation of modern, cloud-native data platforms across AWS and Azure, using Databricks, Microsoft Fabric, and cloud-native services. You will own delivery across the full data lifecycle — ingestion, transformation, modelling, and serving — enabling real-time analytics, reporting, and AI use cases.

This is a hands-on senior role. You will still be in the code every day, but you will also be the person clients and engineers look to for technical direction: shaping the architecture, making the trade-off calls, setting the standards, and raising the level of everyone around you.

Much of our work centres on lakehouse platforms, so proven depth in Databricks and/or Microsoft Fabric is essential.

Key Responsibilities

The list below is broad. We expect production-grade experience across a substantial portion of it, and the judgement to know which tools a given problem actually calls for.

Technical Leadership and Delivery Ownership
  • Act as technical lead on client engagements: own the solution design end to end, from discovery and architecture through to a supported production platform.
  • Translate ambiguous business requirements into pragmatic designs, and communicate them credibly to engineers and senior client stakeholders.
  • Make and justify architectural trade-offs.
  • Set and enforce engineering standards: code review, testing, branching and release strategy, definition of done.
  • Mentor and grow junior and mid-level engineers through pairing, review and feedback; contribute to hiring and technical assessment.
  • Contribute to internal capability: reusable accelerators, reference architectures, and solution-shaping input.
Software Engineering Foundations
  • Strong grounding in software engineering fundamentals (data structures, algorithms, design patterns), and the ability to hold others to that standard.
  • Advanced proficiency in Python and SQL (additional languages advantageous).
  • Deep experience with Git and CI/CD — able to design a branching and deployment strategy for a team, not just work within one.
  • Hands-on Infrastructure as Code with Terraform or Bicep, including the Databricks provider and Fabric deployment automation; able to own a platform's IaC estate.
  • Promoting data assets as code: Databricks Asset Bundles and Repos, Fabric Git integration and deployment pipelines.
  • Comfortable in Linux environments.
Ingestion and Streaming
  • Design scalable ingestion across hybrid and cloud environments, and the standards other engineers will follow.
  • Streaming: AWS Kinesis / MSK (Kafka), Azure Event Hubs, Databricks Structured Streaming, Fabric Eventstream and Real-Time Intelligence (Eventhouse / KQL).
  • Incremental and batch: Auto Loader, Lakeflow Connect, Fabric Dataflows Gen2 and Copy Job, AWS DataSync / DMS, Azure Data Factory / Synapse / Fabric Data Pipelines.
  • Integration via APIs, JDBC/ODBC and CDC, including Fabric Mirroring (Azure SQL, Cosmos DB, Snowflake, Unity Catalog) and Delta Sharing.
Storage and Lakehouse
  • Design and manage data lakes on Amazon S3 and Azure Data Lake Storage Gen2, and architect lakehouse platforms using:
  • Databricks — Delta Lake, Unity Catalog (catalogs, schemas, external locations, storage credentials, volumes), managed vs external tables, streaming tables and materialised views, Lakehouse Federation, Delta Sharing.
  • Microsoft Fabric — OneLake and shortcuts, Lakehouse, Warehouse, Fabric SQL Database, workspace and domain design, capacity (F SKU) sizing and CU management.
  • Table optimisation: partitioning, liquid clustering, Z-ordering, OPTIMIZE / VACUUM, V-Order and statistics, small-file management, time travel and retention.
  • Open formats and standards: Delta Lake, Parquet, Avro, JSON, Apache Iceberg / Delta UniForm.
  • Relational (Postgres, SQL Server, Aurora), NoSQL (DynamoDB, Cosmos DB) and caching (Redis).
Processing and Transformation
  • Build scalable ETL/ELT pipelines using:
  • Databricks — PySpark and Spark SQL, Delta Live Tables / Lakeflow Declarative Pipelines, Workflows, serverless and job compute, Photon, cluster policies, dbt.
  • Microsoft Fabric — Spark notebooks and job definitions, environments and library management, Warehouse T-SQL and stored procedures, Dataflows Gen2.
  • AWS / Azure native — Glue, Lambda, EMR, Synapse Spark.
  • Own the data model: medallion architecture (Bronze/Silver/Gold), dimensional models (star schemas, SCD handling), and platform-wide modelling standards.
  • Build reusable, testable pipelines: modular notebooks and Python packages, parameterisation, unit and integration testing, idempotent and restartable loads.
  • Lead Spark performance and cost tuning: skew and shuffle management, caching, right-sizing compute, autoscaling — and diagnose the problems others cannot.
Analytics and AI Enablement
  • Design platforms that support BI, advanced analytics, machine learning and AI use cases:
  • Databricks — SQL warehouses, AI/BI dashboards and Genie, MLflow, Unity Catalog model registry, Model Serving, Mosaic AI (Vector Search, agent and RAG patterns), feature engineering.
  • Microsoft Fabric — semantic models and Direct Lake, Power BI report and dataset design, DAX, Fabric Data Science, Data Agents and Copilot.
  • AWS — Redshift, Athena, SageMaker and Bedrock where relevant.
  • Partner with data scientists and analysts on feature engineering and real-time decisioning systems.
  • Establish data quality, observability and lineage frameworks: DLT expectations, Unity Catalog lineage and system tables, Fabric monitoring hub, freshness and volume alerting.
Security, Governance and Compliance
  • Design secure, enterprise-grade platforms: AWS IAM, Azure Entra ID, RBAC, managed identities and service principals.
  • Governance: Unity Catalog (namespace design, grants and inheritance, row filters and column masks, attribute-based access, audit logs); Fabric (workspace roles and item permissions, OneLake security, sensitivity labels, domains, Purview); AWS Lake Formation.
  • Networking: VPC / VNets, private endpoints, Direct Connect / ExpressRoute, Databricks secure cluster connectivity and Private Link, Fabric private links and managed VNets.
  • Encryption: KMS / Key Vault / TLS, customer-managed keys, secret scopes.
  • Engage with client security and architecture forums to get designs approved.
Orchestration and Operations
  • Orchestrate with Databricks Workflows, AWS Step Functions / MWAA (Airflow), ADF / Synapse and Fabric Data Pipelines.
  • Monitoring and observability: CloudWatch, Azure Monitor, Log Analytics, Databricks system tables and query history, Fabric Capacity Metrics and monitoring hub.
  • Own the operational posture of what you build: SLAs, alerting, incident response, root cause analysis, and runbooks.
  • Drive reliability, performance and cost optimisation (FinOps): compute policies, serverless vs classic, tagging and chargeback, Fabric capacity smoothing and pause/resume.
Requirements
  • Bachelor's degree in Engineering, Computer Science, or a related field.
  • Typically 7+ years in data or software engineering, with at least 3 years building data platforms on AWS and/or Azure.
  • A track record of delivering production data solutions end-to-end — including at least one platform you can talk through in depth: the constraints, the decisions, and what you'd do differently.
  • Substantial hands-on production experience with at least one lakehouse platform (Databricks or Fabric), working knowledge of the other, and the appetite to build depth.
  • Experience leading a workstream or small team: setting technical direction, reviewing work, and mentoring.
  • Confident client-facing communication. Able to run a design workshop, defend an architecture, and explain a trade-off to a non-technical audience.
  • Certifications advantageous, professional or expert level especially: Databricks (Data Engineer Professional / Associate, Machine Learning Associate); Microsoft (DP-700, DP-600, DP-203, AZ-305); AWS (Data Engineer Associate, Solutions Architect Associate / Professional).
  • Experience with Docker, Kubernetes and CI/CD for data workloads.
  • Strong working knowledge of data security and compliance standards.
  • Strong problem-solving skills, with the judgement to know when a simple solution beats a sophisticated one.
  • Self-organising, able to prioritise across multiple tasks and competing stakeholders.
  • Excellent written and verbal communication, and a collaborative approach with clients and colleagues.
  • Willingness to travel to clients as and when required.
What We Offer
  • A culture of engineering and an environment where ideas are heard, and builders can build.
  • Competitive compensation and bonus structure.
  • A flexible, supportive environment that values diversity, work-life balance and personal growth.
  • Career advancement, mentorship and ongoing development.
  • Cutting-edge technology, engaging client projects, and a talented team around you.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineers
Data Engineers

Blue Pearl HQ • Johannesburg

On-site
ZAR 600,000 - 1,100,000
Data Engineers
Data Engineers

Blue Pearl PTY LTD • Johannesburg

On-site
ZAR 600,000 - 1,200,000
Data Engineers
Data Engineers

Blue Pearl • Johannesburg

On-site
ZAR 900,000 - 1,500,000
Data Engineer (Analytics & Data Platform)
Data Engineer (Analytics & Data Platform)

ATS Client • Cape Town

Hybrid
ZAR 1,200,000 - 1,800,000
Training budget
Flexible working arrangements
Career development
Senior Data Scientist (ML Engineering)
Senior Data Scientist (ML Engineering)

Dariel • Sandton

On-site
ZAR 1,200,000 - 1,800,000
Data Engineers
Data Engineers

Blue Pearl PTY • Johannesburg

On-site
ZAR 800,000 - 1,200,000
None
Azure Data Engineer
Azure Data Engineer

Ariston • Randburg

On-site
ZAR 700,000 - 1,000,000
Senior Data Engineer (Databricks & Cloud Data Platforms)
Senior Data Engineer (Databricks & Cloud Data Platforms)

DVT • Cape Town

On-site
ZAR 900,000 - 1,500,000
Senior Data Fabric Data Engineer
Senior Data Fabric Data Engineer

Belay Talent Solutions • Midrand

Hybrid
ZAR 900,000 - 1,350,000
Senior Data Engineer / Platform Lead
Senior Data Engineer / Platform Lead

ATS Client • Pretoria

On-site
ZAR 900,000 - 1,300,000