A recruitment agency is seeking an Intermediate Data Engineer in Durban, South Africa. The successful candidate will have a strong background in Scala or Java, experience with cloud platforms like AWS, GCP, or Azure, and a solid understanding of data modeling principles. Responsibilities include building scalable data infrastructure, developing APIs for data science outputs, and optimizing cloud costs. The ideal candidate will possess a Master's in a related field and have 2-3 years of relevant experience.
Qualifications
Matric (Grade 12) is required.
2-3 years of relevant work experience is a must.
Knowledge of API development and machine learning deployment.
Responsibilities
Build and scale data infrastructure for real-time data processing.
Create scalable data ingestion and machine learning inference pipelines.
Develop APIs to deliver data science outputs.
Scale production systems for increased demand.
Provide visibility into data platform health.
Automate and handle life-cycle of data processing systems.
Skills
Scala
Java
AWS
GCP
Azure
Data modeling
Relational databases
NoSQL databases
Spark
Hadoop
Kafka
Docker
Kubernetes
Education
Masters in Software Engineering, Data Engineering, Computer Science or related field
Tools
Postgres
ElasticSearch/OpenSearch
Graph databases
Job description
About the job Intermediate Data Engineer
Matric (Grade 12)
Masters in Software Engineering, Data Engineering, Computer Science or related field
2-3 years of relevant work experience
Strong Scala or Java background
Knowledge of AWS, GCP, Azure, or other cloud platform
Understanding of data modeling principles
Ability to work with complex data models
Experience with relational and NoSQL databases (e.g. Postgres, ElasticSearch/OpenSearch, graph databases such as Neptune or neo4j)
Experience with technologies that power analytics (Spark, Hadoop, Kafka, Docker, Kubernetes) or other distributed computing systems
Knowledge of API development and machine learning deployment
Responsibilities:
Build and scale data infrastructure that powers real-time data processing of billions of records in a streaming architecture
Build scalable data ingestion and machine learning inference pipelines
Build general-purpose APIs to deliver data science outputs to multiple business units
Scale up production systems to handle increased demand from new products, features, and users
Provide visibility into the health of our data platform (comprehensive view of data flow, resources usage, data lineage, etc) and optimize cloud costs
Automate and handle the life-cycle of the systems and platforms that process our data