Cloud and Data Engineer

Jobtailor

KwaZulu-Natal

Hybrid

ZAR 700,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a data engineer to optimize data pipelines, implement ETL improvements, and collaborate with data scientists to productionize models. You will design AI-enabled data workflows, ensure data quality, and contribute to dashboards and reports.

The role requires strong Python/SQL skills, Databricks expertise, and experience with Delta Lake, DLT, and testing frameworks. Collaboration across teams and CI/CD integration are essential.

Qualifications

  • 3+ years in cloud/data engineering or SaaS environment.
  • Hands-on Databricks experience (notebooks, Delta Lake, DLT pipelines, jobs).
  • Experience with AI/ML pipelines or agentic data workflows is beneficial.

Responsibilities

  • Optimize and enhance data pipelines, ETL processes for performance and reliability.
  • Implement data processing best practices for storage and retrieval.
  • Write unit tests for notebooks and PySpark transformations; integrate with CI/CD.

Skills

Databricks Expertise
Automated Testing
Data Processing Frameworks
Cloud Engineering
AI/ML Pipeline Experience
SQL
Python
Delta Lake
NoSQL Databases
Git
Delta Live Tables
Great Expectations
Semantic Kernel
LangChain
Azure OpenAI

Education

Bachelor's degree in Computer Science or related field

Tools

Delta Live Tables
Databricks
Great Expectations
Apache Spark
Hadoop
Kafka
Git

Job description

  • Optimize and enhance existing data pipelines, ETL processes, and data workflows to improve performance, scalability and reliability.
  • Implement best practices for data processing, storage, and retrieval.
  • Implement automated testing within Databricks using frameworks such as pytest and nutter: write unit tests for notebooks and PySpark transformations, integration tests across pipeline stages, and data quality assertions, with results surfaced through CI/CD pipelines.
  • Design, implement and maintain algorithms to solve technical problems related to data processing and analytics.
  • Collaborate with data scientists to support and optimize machine learning algorithms and models in production.
  • Assist to benchmark and ensure efficient use of computational resources for data processing and algorithm execution.
  • Design and implement AI agentic workflows that automate multi-step data tasks: orchestrate LLM-powered agents to perform data discovery, anomaly investigation, and automated root-cause analysis within the data platform.
  • Apply prompt engineering and model evaluation practices when integrating large language models into data pipelines; enforce responsible-AI guardrails including output validation and human-in-the-loop review for high-impact decisions.
  • Design, implement and maintain solutions for data integration from various sources, ensuring data consistency and integrity.
  • Work on data ingestion and transformation processes to support analytics and reporting needs.
  • Validate integrated data using Databricks-native testing approaches: apply Delta Live Tables expectations, Great Expectations, or equivalent data quality frameworks to enforce schema, completeness, and accuracy contracts at ingestion.
  • Work closely with the technical product owner to understand business requirements and translate them into technical solutions.
  • Provide technical support and guidance to other team members regarding data-related issues.
  • Work closely with cross-functional teams to understand data requirements and deliver solutions that meet business needs.
  • Help to document key processes and services for the purposes of sharing the load and approach to technical support related issues.
  • Work in conjunction with the BI Engineering to develop and maintain both product and internal dashboards and reports.
  • Interpret data to provide meaningful insights and recommendations to stakeholders.
  • Work closely with business teams to understand their and customer reporting needs and deliver tailored solutions.
  • Ensure data accuracy and integrity in all reports and dashboards.
Requirements
  • 3+ years of experience in cloud engineering, data engineering, or a similar role within a SaaS environment.
  • Demonstrable experience with Databricks: notebooks, Delta Lake, DLT pipelines, Jobs, and writing automated tests for data transformations within the platform.
  • Hands-on exposure to AI/ML pipelines or agentic data workflows; experience with Azure OpenAI, Azure AI Search, or equivalent services advantageous.
  • Strong Mathematical, Analytical, Conceptual and Problem-Solving Abilities.
  • Solution Driven.
  • Ability to find the root cause of problems and quickly determine effective solutions.
  • Ability to anticipate risk.
  • Troubleshooting, analytical and attention to details.
  • Ability to prioritize and manage time effectively.
  • Excellent Communication Skills.
  • Proficiency in cloud platforms such as Azure, AWS or Google Cloud.
  • Strong experience with data processing frameworks and tools (e.g., Databricks, Apache Spark, Hadoop, Kafka).
  • Strong Expertise in SQL and experience with NoSQL databases.
  • Familiarity with Git.
  • Proficiency in programming languages such as Python.
  • Experience building AI agentic systems: multi-agent orchestration, tool/function calling, RAG pipelines, or similar autonomous workflow patterns applied to data engineering problems.
  • Working knowledge of AI/LLM frameworks relevant to data engineering such as Semantic Kernel, LangChain, AutoGen, or the Azure OpenAI Service SDK; familiarity with prompt engineering and model evaluation.
  • Familiarity with vector databases and embedding models (e.g. Azure AI Search, Chroma, pgvector) advantageous; understanding of retrieval-augmented generation (RAG) patterns a plus.
  • Proficiency in automated testing within Databricks: pytest, nutter, Delta Live Tables expectations, or Great Expectations; ability to integrate test runs into Azure DevOps or equivalent CI/CD pipelines.
Core Competencies

Demonstrates expertise in optimizing data pipelines and ETL processes, with a strong focus on automated testing and data quality assurance within Databricks. Proficient in collaborating with cross-functional teams to deliver data-driven solutions that meet business needs while ensuring data integrity and accuracy.

Highest-signal resume keywords
  • Databricks Expertise
  • Automated Testing with Pytest
  • Data Processing Frameworks
  • Cloud Engineering Proficiency
  • AI/ML Pipeline Experience
ATS Optimization Keywords
Hard Skills
  • SQL
  • Python
  • Databricks
  • Apache Spark
  • Hadoop
  • Kafka
  • Delta Lake
  • NoSQL Databases
  • Automated Testing
  • Data Integration
Soft Skills
  • Problem-Solving Abilities
  • Attention to Detail
  • Excellent Communication Skills
  • Time Management
  • Solution Driven
Industry Keywords
  • Data Engineering
  • SaaS Environment
  • AI Agentic Workflows
  • Data Quality Frameworks
  • CI/CD Pipelines
Tools & Technologies
  • Azure
  • AWS
  • Google Cloud
  • Azure OpenAI
  • Azure AI Search
  • Git
  • Delta Live Tables
  • Great Expectations
  • Semantic Kernel
  • LangChain
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data/AI Architect (Databricks)
Data/AI Architect (Databricks)

Blue Pearl PTY • Johannesburg

On-site
ZAR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

cloudandthings.io • Wes-Kaap

On-site
ZAR 720,000 - 1,100,000
Competitive compensation
Flexible and supportive work culture
Career advancement opportunities
+1
Data Engineer (MLOps / Analytics Focus)
Data Engineer (MLOps / Analytics Focus)

Blue Pearl HQ • Johannesburg

On-site
ZAR 720,000 - 900,000
Lead Data Engineer/Senior Data Engineer
Lead Data Engineer/Senior Data Engineer

Indsafri • South Africa

On-site
ZAR 800,000 - 1,200,000
Data Engineer
Data Engineer

cloudandthings.io • City of Johannesburg Metropolitan Municipality

On-site
ZAR 600,000 - 1,200,000
Competitive compensation
Flexible work environment
Career development
+1
Senior Data Engineer - GenAI, Databricks & Lakehouse
Senior Data Engineer - GenAI, Databricks & Lakehouse

Indsafri • South Africa

On-site
ZAR 800,000 - 1,200,000
Data Engineering Lead
Data Engineering Lead

Blue Pearl HQ • Johannesburg

On-site
ZAR 1,100,000 - 1,900,000
Senior Software Developer
Senior Software Developer

Jobtailor • KwaZulu-Natal

On-site
ZAR 900,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

BETSoftware • Johannesburg

On-site
ZAR 800,000 - 1,200,000
Data Engineer (MLOps / Analytics Focus)
Data Engineer (MLOps / Analytics Focus)

Blue Pearl PTY • Johannesburg

On-site