Data Engineer

DataProphet

Cape Town

On-site

ZAR 720,000 - 1,100,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DataProphet seeks a Data Engineer to design, build and maintain data infrastructure and pipelines enabling AI solutions. You will handle diverse data sources to ensure reliable, accessible data for scientists and ML systems.

You will own data pipelines end-to-end from ingestion to delivery, build clean datasets, and contribute to data models and lakehouse structures, with a focus on quality, monitoring and operational excellence.

Qualifications

  • Production data infrastructure experience required.
  • End-to-end ownership of data pipelines from ingestion to delivery.
  • Experience with data quality, monitoring and incident response.
  • Certifications in cloud or data platforms are a plus.

Responsibilities

  • Design, build, test and maintain scalable production data pipelines & infrastructure.
  • Own data pipelines end-to-end across ingestion, transformation, storage and delivery.
  • Integrate data from varied source systems into DataProphet's data environment.
  • Build reliable, clean and usable datasets for Data Scientists and ML systems.
  • Design and maintain data models and data warehouse/lakehouse structures.
  • Implement data quality validation, monitoring and alerting.
  • Diagnose and resolve pipeline failures and data-quality issues.
  • Build prototypes and PoC solutions as needed.
  • Evaluate hardware, software and cloud solutions for data systems.
  • Contribute to data and system architecture and design.
  • Integrate new data management tools and create analytics components.
  • Install and update disaster recovery procedures.

Skills

SQL
Python
Data modelling
Data warehousing
Airflow
Apache Spark
Git
Databricks
Snowflake
Terraform
Kafka
Kinesis
Cloud platforms
CI/CD
System design

Education

Bachelor's / Honours / Master’s / PhD in Computer Science or related field

Tools

Airflow
Dagster
Apache Spark
Prefect
Git
Databricks
ClickHouse
Snowflake
DuckDB
Terraform

Job description

Introduction:

DataProphet is a global leader in Artificial Intelligence (AI) for manufacturing. Our award winning technology embeds unique adaptations and advancements of deep learning, enabling AI to have a significant, practical, impact on the factory floor. DataProphet’s solutions are built to be adapted and integrated into existing environments, making it possible for our digital transformation team to take your operations from zero to AI. We understand manufacturing and that real impact is achieved with pre-emptive actions because real-time is often too late. For more information, visit www.dataprophet.com Why join DataProphet?

You’ll work on technically challenging problems where AI moves beyond experimentation and creates measurable real-world impact.

You’ll have meaningful ownership, work alongside highly capable colleagues across Data Science and Engineering, and have the opportunity to apply your skills across new problems, use cases and domains.

Curiosity, continuous learning and collaboration are central to how we work. Our team works together from our DeWaterkant, Cape Town office in a professional, supportive environment designed to help people do their best work.

Role Overview:

We are looking for a Data Engineer to design, build and maintain the data infrastructure and pipelines that enable DataProphet's AI solutions.

You will work with complex and varied data sources and be responsible for making data reliable, accessible and usable by Data Scientists, machine learning systems and other downstream consumers.

Roles and responsibilities will include, but are not limited to:

  • Design, build, test and maintain scalable production data pipelines & infrastructure..
  • Own data pipelines end-to-end across ingestion, transformation, storage and delivery.
  • Integrate data from varied source systems into DataProphet's data environment.
  • Build reliable, clean and usable datasets for Data Scientists, machine learning systems and other stakeholders.
  • Design and maintain appropriate data models and data warehouse / lakehouse structures.
  • Implement data quality validation, monitoring and alerting.
  • Diagnose and resolve pipeline failures and data-quality issues.
  • Build prototypes and proof-of-concept solutions where required.
  • Evaluate hardware, software and cloud solutions for building and integrating systems & data warehouses.
  • Contribute to data & system architecture and design.
  • Integrate new data management technologies and software engineering tools into existing structures and create custom software components and analytics applications.
  • Install and update disaster recovery procedures.
Qualifications & Experience:
  • Bachelor's / Honours / Master’s / PhD Computer Science, Software Engineering, or a related field.
  • 2–5 years of experience building and maintaining production data infrastructure.
  • Track record of owning features end-to-end: design, implementation, testing, deployment, and post-release support
  • Demonstrated experience owning data pipelines from ingestion through transformation, storage and downstream delivery.
  • Experience working with data quality, monitoring and incident response.
  • Relevant technical or cloud certifications are advantageous.
Core skills
  • Strong SQL and Python skills, with the ability to write production-grade, maintainable and testable code.
  • Strong understanding of data modelling, data warehousing and modern data architecture.
  • Experience building and orchestrating production data pipelines using tools such as Airflow, Dagster, Apache Spark or Prefect.
  • Experience with version control tools such as Git and collaborative development
  • Familiarity with modern data platforms like Databricks, ClickHouse, Snowflake, or DuckDB .
  • Experience with data transformation frameworks.
  • Experience working with at least one major cloud platform — AWS, Azure or GCP.
  • Understanding of data quality, testing, monitoring and observability.
  • Familiarity with CI/CD practices for data pipelines as well as production concerns (logging, monitoring, performance, security)
  • Ability to translate loosely defined data requirements into effective data engineering solutions.
  • Exposure to distributed processing technologies such as Apache Spark, streaming technologies such as Kafka or Kinesis, and infrastructure-as-code tools such as Terraform.
  • Familiarity with networking performance, design and tradeoffs.
  • Comfort reading and adapting to unfamiliar codebases.
  • Solid grasp of system design at a component level (can design a service or module, not just a function).
  • Can break down a loosely defined problems into a technical plans.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

DataProphet • Cape Town

On-site
ZAR 700,000 - 1,000,000
Data Scientist
Data Scientist

DataProphet • Cape Town

On-site
ZAR 600,000 - 900,000
Data Engineer: Build Scalable AI Data Pipelines
Data Engineer: Build Scalable AI Data Pipelines

DataProphet • Cape Town

On-site
ZAR 720,000 - 1,100,000
Data Engineer
Data Engineer

Network Recruitment • Cape Town

On-site
ZAR 700,000 - 1,100,000
Lead Data Engineer
Lead Data Engineer

Key Recruitment Group • Wes-Kaap

On-site
ZAR 1,000,000 - 1,400,000
Senior Data Solutions Engineer
Senior Data Solutions Engineer

Future Fit • Johannesburg

Hybrid
ZAR 900,000 - 1,500,000
Competitive compensation package
Twice-yearly salary increases
Employee wellness programs
+1
Senior Data Engineer
Senior Data Engineer

Future Fit • Johannesburg

On-site
ZAR 900,000 - 1,300,000
Data Engineer
Data Engineer

Hire Resolve • Sandton

On-site
ZAR 700,000 - 1,000,000
Data Engineering Lead
Data Engineering Lead

Blue Pearl PTY • Johannesburg

On-site
ZAR 1,200,000 - 1,900,000
Data Engineer
Data Engineer

PBT Group • Cape Town

On-site
ZAR 900,000 - 1,300,000