Client-Facing Data Engineer — Etl, Snowflake & Python Expert

Communicate Finance

South Africa

Hybrid

ZAR 600,000 - 900,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Hybrid work model

Job summary

Communicate Finance in South Africa is seeking an experienced Data Engineer in Soweto to design, build, and maintain scalable data pipelines for AI/ML workloads. You will develop ETL/ELT processes to ingest, transform, and load data from APIs, databases, and other sources, while ensuring data quality, security, and governance.

The role requires a strong background in data engineering, big data technologies, and cloud data services, with emphasis on PII handling and regulatory compliance.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related quantitative field.
  • 3+ years of experience in data engineering, with big data focus.
  • Proficiency in Python and SQL, and data processing concepts.

Responsibilities

  • Design, build, and maintain scalable data pipelines for AI/ML workloads.
  • Develop ETL/ELT processes to ingest, transform, and load data from diverse sources.
  • Ensure data quality, integrity, and security across platforms.

Skills

Python
SQL
Data Warehouse
Databases
Cloud
Analytics
Data modelling
ETL/ELT
Data Pipeline Development
CI/CD

Education

Bachelor's degree in Computer Science or Engineering

Tools

Spark
Hadoop
Azure/AWS/GCP
Microsoft Fabric

Job description

An established insurance company is seeking to hire a highly skilled and experienced Data Engineer to join their team. Your:

  • Formal Education:
    • Degree in Data Science, Information Technology, Computer Science or equivalent
  • Advantageous:
    • Cloud Data Certifications
    • Exposure to regulated environments,financial services, fintech
  • Experience:
    • Minimum of 2 years in a data engineer role or a similar technical role.
  • Responsibilities:
    • Build and maintain ETL pipelines supporting a multi-tenant data platform, ingesting data from APIs, databases, and event sources.
    • Build and maintain Data Platform APIs that allow teams to ingest, process, and access data easily and reliably.
    • Implemented tenant-specific logic by following existing configuration and naming conventions.
    • Apply tenant‑level data isolation using schemas, partitions, or access controls.
    • Build models from existing templates used for financial and operational reporting. Develop models for analytics and reporting, maintaining consistency with shared data models.
    • Monitor scheduled pipelines, investigate failures, and resolve data quality issues and inconsistencies.
    • Maintain daily and incremental data loads into the data warehouse.
    • Assist with onboarding new clients by validating source data and testing pipeline outputs.
    • Work closely with senior data engineers to learn patterns for multi‑tenant data isolation.
    • Collaborate with analytics, product, and customer facing teams to understand reporting needs.
    • Support strict regulatory and audit requirements by following data handling,retention, and audit guidelines.
    • Handle financial and sensitive data (PII) according to company policies and regulatory standards (e.g. POPIA)
    • Apply least‑privilege access and rolebased access controls, and support data protection through masking, encryption,and established security standards.
  • Technical Skills:
    • Programming languages Good knowledge of programming languages such as Python, especially used for pipeline and data manipulation.
    • SQL working experience using SQL for data cleaning, aggregation, data transformation and integration.
    • Data Warehouse - have a fundamental understanding of data warehousing solutions and platforms.
    • Databases hands on experience working with relational and non‑relational databases.
    • Cloud computing comfortable building data solutions using cloud hosted services or data platforms.
    • Analytics skills strong problem solving skills, understand data characteristics, identify patterns, and data quality issues.
    • Data modelling and ETL Is able to communicate and translate business requirements into existing data models.
    • Data Pipeline Development build and validate smaller scale data pipelines independently.
    • CI/CD and Version Control apply best practice for managing pipelines and data workflows.
  • Our client is seeking a meticulous and experienced AI Data Engineer to join their team in Soweto. This role is critical for building and maintaining the robust data infrastructure that powers our AI and machine learning initiatives. You will be responsible for designing, developing, and optimizing data pipelines, ensuring data quality, availability, and accessibility for data scientists and ML engineers. The ideal candidate possesses a strong background in data engineering principles, experience with big data technologies, and a passion for enabling data-driven innovation through reliable and efficient data systems.
  • About the Role
    • Our client is seeking a meticulous and experienced AI Data Engineer to join their team in Soweto. This role is critical for building and maintaining the robust data infrastructure that powers our AI and machine learning initiatives. You will be responsible for designing, developing, and optimizing data pipelines, ensuring data quality, availability, and accessibility for data scientists and ML engineers. The ideal candidate possesses a strong background in data engineering principles, experience with big data technologies, and a passion for enabling data-driven innovation through reliable and efficient data systems.
  • Key Responsibilities
    • Design, build, and maintain scalable and reliable data pipelines for AI/ML workloads.
    • Develop ETL/ELT processes to ingest, transform, and load data from various sources.
    • Ensure data quality, integrity, and security across all data platforms.
    • Collaborate with data scientists and ML engineers to understand data requirements and provide solutions.
    • Optimize data infrastructure for performance, cost, and scalability.
    • Implement data governance policies and best practices.
  • Requirements
    • Bachelor's degree in Computer Science, Engineering, or a related quantitative field.
    • 3+ years of experience in data engineering, with a focus on big data technologies.
    • Proficiency in programming languages such as Python, SQL, or Scala.
    • Experience with big data platforms like Spark, Hadoop, or similar.
    • Familiarity with cloud data services (AWS, Azure, GCP) and data warehousing concepts.
    • Strong understanding of data modeling, database design, and data architecture.
  • Benefits
    • Competitive salary and comprehensive benefits package.
    • Opportunities for professional growth and training in AI and data technologies.
    • Access to modern data infrastructure and cutting‑edge tools.
    • A collaborative and innovative work environment in Soweto.
    • Hybrid work model offering flexibility.
  • Bachelor's degree in Computer Science, Information Systems, Data Engineering or a related field.
  • Relevant Microsoft Fabric and/or Azure data certification.
  • 0-1 years of practical data engineering experience, including strong recent hands‑on experience with Microsoft Fabric.
  • Strong hands‑on experience with Microsoft Fabric, including OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
  • Strong SQL Server and T‑SQL capability, including complex query development, schema design, indexing and performance optimisation.
  • Practical experience developing, maintaining and supporting production ETL/ELT pipelines.
  • Experience integrating and extracting data from REST/SOAP APIs, databases, flat files and other structured or unstructured data sources.
  • Proficiency in data transformation using SQL and Python/PySpark, with an understanding of scalable data processing practices.
  • Practical experience in data warehousing, dimensional modelling, incremental loading, orchestration and schema evolution.
  • Experience with troubleshooting pipeline failures, data quality issues and performance bottlenecks, including the ability to restore service efficiently.
  • Experience with source control, CI/CD and deployment practices using Git, Azure DevOps or equivalent tools.
  • Experience supporting Power BI and other downstream analytical or reporting requirements.
  • Demonstrated ability to take ownership of an existing technical environment with limited hand‑handing.
  • Strong documentation, communication and stakeholder engagement skills.
  • Microsoft Fabric: OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
  • SQL Server / T‑SQL
  • Python / PySpark
  • REST/SOAP APIs and structured/unstructured data ingestion.
  • ETL/ELT, incremental loading, orchestration and scheduling.
  • Dimensional modelling, medallion architecture, schema evolution and data warehousing.
  • Power BI integration and understanding of downstream analytical requirements.
  • Git / Azure DevOps, CI/CD and environment deployment practices.
  • Monitoring, data quality, performance optimisation, security and operational support.
  • Proficient in Afrikaans and English.
  • Own transport and valid drivers license.
  • Take ownership of the existing Microsoft Fabric data engineering environment and become productive quickly following handover.
  • Design, develop, maintain and orchestrate reliable batch and near‑real‑time data ingestion pipelines using appropriate Microsoft Fabric capabilities.
  • Extract and ingest data from structured and unstructured sources, including REST APIs, SOAP APIs, databases and flat files.
  • Develop robust data transformation logic using SQL, Python/PySpark, Fabric notebooks and Dataflows Gen2, as appropriate.
  • Implement incremental loading, retry mechanisms, logging, monitoring and alerting to support data integrity and pipeline reliability.
  • Troubleshoot and resolve pipeline failures and data processing issues efficiently.
  • Optimise data pipelines and processing workloads for performance, scalability and cost‑effectiveness.
  • Design, manage and evolve scalable data architectures using Microsoft Fabric, OneLake, Lakehouse, Warehouse and SQL Server.
  • Maintain appropriate data‑layering and medallion architecture principles, where applicable, with clear movement from raw to curated data.
  • Develop and maintain robust schema designs, indexes, partitioning and query strategies to support analytical and operational workloads.
  • Manage schema evolution and version control to maintain consistency and minimise disruption to downstream consumers.
  • Maintain metadata, data dictionaries, architecture documentation and technical documentation to improve supportability and reduce key‑person dependency.
  • Define and maintain appropriate role‑based access and security controls.
  • Build and maintain analytical data stores using Microsoft Fabric Warehouse and/or Lakehouse patterns.
  • Apply appropriate data‑loading, partitioning, storage optimisation and query‑performance practices.
  • Develop and maintain stable, well‑modelled datasets for Power BI and other analytical consumers.
  • Work with reporting and analytical teams to investigate and resolve data‑related issues.
  • Ensure data structures and outputs support downstream reporting and business intelligence requirements.
  • Develop and maintain conceptual, logical and physical data models.
  • Apply dimensional modelling techniques, including star and snowflake schemas, to support analytics and reporting.
  • Apply appropriate normalisation and relational modelling techniques for operational and analytical workloads.
  • Ensure consistency of data models across systems.
  • Manage schema versioning and evolution without unnecessarily disrupting downstream consumers.
  • Apply agreed data engineering standards and modelling principles consistently.
  • Work independently and take end‑to‑end ownership of assigned data engineering deliverables, incidents and production issues.
  • Provide clear and timely updates regarding progress, risks, dependencies and blockers.
  • Engage directly with technical and business stakeholders to clarify requirements and agree practical solutions.
  • Explain technical concepts and trade‑offs in a manner appropriate to the relevant stakeholder.
  • Maintain practical technical documentation, including runbooks, architecture notes, change logs and release notes.
  • Take accountability for the successful delivery and operational support of assigned solutions.
  • Automate recurring data engineering and operational activities where practical.
  • Implement monitoring and alerting to identify data quality issues, pipeline failures and abnormal processing behaviour.
  • Analyse and optimise query, notebook and pipeline performance across SQL Server and Microsoft Fabric.
  • Monitor capacity and resource utilisation and contribute to scalability and cost‑control decisions.
  • Deploy solutions using appropriate CI/CD and controlled deployment practices.
  • Apply data security best practices, including secure authentication, least‑privilege access and appropriate encryption.
  • Ensure data engineering solutions comply with applicable data governance policies and regulatory requirements.
  • Apply sound engineering practices relating to recoverability, auditability, supportability and controlled change.
  • Protect confidential and sensitive business information.
  • Collaborate with developers, data analysts, data scientists and business stakeholders to understand requirements and deliver practical solutions.
  • Support effective handover and knowledge transfer to reduce key‑person dependency within the data environment.
  • Share technical knowledge and contribute to continuous improvement of team practices and the data environment.
  • Provide guidance and support to junior team members where required.
  • Remain accountable for the quality, reliability and timeliness of own deliverables.
  • Document data processes, transformations, dependencies and architectural decisions.
  • Validate data outputs through reconciliation, data quality checks and appropriate testing before production deployment.
  • Maintain high standards of engineering quality by following agreed development, code review, testing, deployment, backup and archival practices.
  • Ensure changes are appropriately tested, documented and controlled before implementation.
  • Safeguard confidential information and data.
  • Support compliance with applicable organisational policies, standards and regulatory requirements.
  • Market related Remuneration Offered
  • Perform exploratory data analysis (EDA) and validate datasets.
  • Use Python extensively for data analysis, investigation and problem‑solving.
  • Work with real‑world and imperfect datasets to identify patterns, issues and insights.
  • Support analytics, reporting and insight‑driven initiatives.
  • Translate client and business questions into clear data outputs and findings.
  • Apply data engineering best practices to ensure data is reliable and fit for analytical use.
  • Investigate data and communicate findings clearly to both technical and business stakeholders.
  • Engage with stakeholders to understand analytical requirements and provide data‑driven solutions.
  • Key Responsibilities:
    • Design, develop and maintain scalable data pipelines for AI and machine learning projects.
    • Implement and manage data warehousing solutions and data lakes, ensuring data integrity and accessibility.
    • Collaborate with data scientists and ML engineers to understand data requirements and deliver robust data solutions.
    • Develop and maintain data governance policies and ensure compliance with data privacy regulations.
    • Monitor data infrastructure performance, troubleshoot issues, and implement improvements for efficiency and reliability.
  • Requirements:
    • Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field.
    • 5+ years of experience in data engineering, with a focus on supporting AI/ML workloads.
    • Strong proficiency in SQL, Python, and data processing frameworks (e.g., Spark, Hadoop).
    • Experience with cloud data services (AWS Redshift, Azure Data Lake, Google BigQuery).
    • Solid understanding of ETL/ELT processes and data modelling techniques.
  • Benefits:
    • Competitive salary and comprehensive health and retirement benefits.
    • Opportunity to work on impactful AI and data initiatives.
    • Professional development and continuous learning opportunities.
    • Hybrid work model providing flexibility.
    • Dynamic team environment in the heart of Sandton.
  • Benefits:
    • Competitive salary and performance-based bonuses.
    • Comprehensive medical aid and retirement benefits.
    • Opportunities for professional development and training in cutting‑edge data technologies.
    • Hybrid work model offering flexibility and team collaboration.
    • Contribute to critical data infrastructure that powers advanced AI capabilities.
  • Benefits:
    • Competitive salary and comprehensive benefits package.
    • Opportunity to work on impactful AI and data initiatives.
    • Professional development and continuous learning opportunities.
    • Hybrid work model providing flexibility.
    • Dynamic team environment in the heart of Sandton.
  • Receive personalized job recommendations.
  • Please wait while we submit your request.
  • You will receive the job details via email shortly.
  • Get your free, confidential resume review.

    or drag and drop your file here.

    Similar jobs

    Similar jobs worth comparing

    Microsoft Fabric Specialist / Data Engineer
    Microsoft Fabric Specialist / Data Engineer

    Praesignis • South Africa

    Hybrid
    ZAR 900,000 - 1,300,000
    Hybrid work model
    Professional development
    Medical aid
    +2
    Senior Ai Data Engineer - Hybrid Pipelines For Ml
    Senior Ai Data Engineer - Hybrid Pipelines For Ml

    Placements24 • Gauteng

    Hybrid
    ZAR 972,000 - 1,188,000
    Hybrid work model
    Competitive salary
    Career growth opportunities
    Data and AI Engineer
    Data and AI Engineer

    WatersEdge Solutions • Johannesburg

    Hybrid
    ZAR 600,000 - 900,000
    Data and AI Engineer
    Data and AI Engineer

    WatersEdge Solutions • Gauteng

    Hybrid
    ZAR 700,000 - 1,100,000
    Intermediate Data Engineer
    Intermediate Data Engineer

    Cls • Pretoria

    On-site
    ZAR 700,000 - 900,000
    Data Engineer (Analytics & Data Platform)
    Data Engineer (Analytics & Data Platform)

    ATS Client • Cape Town

    Hybrid
    ZAR 1,200,000 - 1,800,000
    Training budget
    Flexible working arrangements
    Career development
    Senior Data Engineer - AI Focus
    Senior Data Engineer - AI Focus

    Placements24 • Sandton

    Hybrid
    ZAR 1,200,000 - 1,800,000
    Salary bonuses
    Hybrid work model
    Health and retirement benefits
    +2
    Senior AI Data Engineer
    Senior AI Data Engineer

    Placements24 • Sandton

    Hybrid
    ZAR 900,000 - 1,300,000
    Hybrid work model
    Competitive salary
    Professional development
    +1
    Data Engineer: Fintech Multi-Tenant Etl & Analytics
    Data Engineer: Fintech Multi-Tenant Etl & Analytics

    Competent Candidates • Gauteng

    Hybrid
    ZAR 720,000 - 980,000
    Competitive salary
    Hybrid work model
    Growth opportunities
    Data Engineer (Fmcg / Food / Retail Industry Experience)
    Data Engineer (Fmcg / Food / Retail Industry Experience)

    Merand Recruitment • South Africa

    On-site
    ZAR 600,000 - 900,000