A complete application in a minute — tailored resume and cover letter, ready to send.
Blue Pearl in Johannesburg seeks an experienced Data Engineer to design scalable data platforms and modernise cloud-native data ecosystems for enterprise clients. You will lead end-to-end data pipelines using Python/SQL and tools like Azure Data Factory, AWS Glue, Google Dataflow, Databricks and dbt.
The role requires strong cloud experience across Azure, AWS or GCP, with a focus on CI/CD, data governance and stakeholder collaboration to deliver robust analytics and insights.
Design and build scalable data platforms using modern cloud-native and Lakehouse architectures
Develop and optimise data pipelines using Python, SQL, and tools such as Azure Data Factory, AWS Glue, Google Cloud Dataflow, Databricks, and dbt
Modernise legacy data environments, migrating from on-premises solutions to cloud-native platforms such as Microsoft Fabric, Azure Synapse Analytics, AWS Redshift, Google BigQuery, or Databricks
Engage with clients to conceptualize data solutions aligned to their business strategy
Support our sales team with pre-sales activities, proof-of-concept deliveries, and technical proposals
Provide technical guidance and mentorship to junior and intermediate consultants
Lead technical reviews and contribute to consultants' growth plans
Identify opportunities to automate manual processes, optimise data delivery, and improve infrastructure scalability
Work with stakeholders, including executive, product, and analytics teams, to address data infrastructure needs
Drive knowledge sharing through technical blogs, internal forums, and workshops
Balance billable project work with team support responsibilities
3-5 years' experience
3-5 years of hands-on experience in data engineering.
Strong proficiency in Python and/or SQL , including query optimisation.
Experience working with both relational and non-relational databases.
Experience designing and building data pipelines and data models.
Understanding and practical experience with lakehouse architectures , including the medallion pattern.
Practical experience with at least one major cloud platform, including:
Microsoft Azure
AWS
Google Cloud Platform (GCP)
Familiarity with:
Databricks
Snowflake
Delta Lake
PySpark
Understanding of data transformation frameworks such as dbt .
Experience with version control using Git .
Understanding of CI/CD practices for data workflows.
Strong analytical and problem-solving skills.
Ability to perform root-cause analysis on complex data issues.
Good communication and stakeholder engagement skills.
6-8+ years' experience
6-8+ years of hands-on experience in data engineering.
All intermediate-level technical requirements, together with demonstrable experience in:
Leading end-to-end data platform delivery.
Architecting enterprise-grade lakehouse environments.
Implementing data mesh patterns.
Infrastructure-as-code using tools such as Terraform, Bicep, AWS CDK or Pulumi.
DevOps and CI/CD pipelines.
Working effectively with cross-functional teams in a dynamic consulting environment.
Mentoring junior engineers.
Contributing to technical strategy and solution direction.
Bachelor's degree in:
Computer Science
Information Systems
Information Technology
or a related field.
Master's degree in a relevant field is advantageous.
One or more of the following certifications would be advantageous:
Microsoft Fabric Data Engineer Associate
Microsoft Azure Data Engineer Associate
Databricks Certified Data Engineer Associate
Google Professional Data Engineer
AWS Certified Data Engineer - Associate
Databricks Certified Data Engineer Professional
Python
PySpark
SQL
dbt
Microsoft Fabric Lakehouses
Fabric Pipelines
Fabric Semantic Models
Direct Lake
Azure Data Factory
Azure Data Lake Storage Gen2
Azure Synapse Analytics
Azure Databricks
Azure Event Hubs
BigQuery
Cloud Storage
Dataflow
Dataproc
Pub/Sub
Amazon S3
AWS Glue
Amazon Redshift
Amazon EMR
Amazon Kinesis
Databricks
Delta Lake
Unity Catalog
MLflow
Databricks Workflows
Azure SQL
Azure Cosmos DB
PostgreSQL
Snowflake
BigQuery
Amazon Redshift
Git
Azure DevOps
GitHub Actions
Terraform
Bicep
AWS CDK
CI/CD pipelines
Azure Event Hubs
Azure Stream Analytics
Apache Kafka
Amazon Kinesis
Google Pub/Sub
Microsoft Power BI
Microsoft Fabric Real-Time Dashboards
Looker / Looker Studio
Amazon QuickSight