A technology solutions firm in Metro Manila seeks an experienced data engineer to implement data ingestion processes and design solutions on Cloud platforms. The ideal candidate has extensive knowledge of data integration technologies and streaming frameworks like Kafka, along with a strong background in data analytics. This role involves working with both structured and unstructured data, requiring a bachelor's degree in a relevant field and at least 3 years of industry experience.
Qualifications
3+ years of experience integrating data into analytical platforms.
Proficient in streaming frameworks like Kafka.
Familiarity with data profiling and SQL optimization.
Responsibilities
Implement requests for ingestion and preparation of data sources.
Design solutions on On-premises and Cloud platforms.
Profile source data and validate it for use.
Skills
Data ingestion technologies
Python
Cloud platforms
SQL optimization
Streaming frameworks (Kafka/Spark Streaming)
ETL development
Education
Bachelor’s degree in relevant field
Tools
Hadoop
Jenkins
Bitbucket
Cloud storage solutions
Job description
Responsibilities
Implement request for ingestion, creation, and preparation of data sources
Develop and execute jobs to import data periodically/ (near) real-time from an external source
Setup a streaming data source to ingest data into the platform
Delivers data sourcing approach and data sets for analysis, with activities including data staging, ETL, data quality, and archiving
Design a solution architecture on both On-premises and Cloud platforms to meet business, technical and user requirements
Profile source data and validate fit-for-purpose
Works with Delivery lead and Solution Architect to agree pragmatic means of data provision to support use cases
Understands and documents end user usage models and requirements
Qualifications
Bachelor’s degree in maths, statistics computer science, information management, finance or economics
At least 3 years’ experience integrating data into analytical platforms using patterns like API, files, XML, json, flatfiles, Hadoop file formats, and Cloud file formats.
Experience in ingestion technologies (e.g. sqoop, nifi, flume), processing technologies (Spark/Scala) and storage (e.g. HDFS, HBase, Hive) are essential
Experience in designing and building data pipelines using Cloud platform solutions and native tools.
Experience in Python, JVM-compatible languages, use of CICD tools like Jenkins, Bitbucket, Nexus, Sonarqube
Experience in data profiling, source-target mappings, ETL development, SQL optimisation, testing and implementation.
Expertise in streaming frameworks (Kafka/Spark Streaming/Storm) essential
Experience managing structured and unstructured data types
Experience in requirements engineering, solution architecture, design, and development / deployment
Experience in creating big data or analytics IT solution
Track record of implementing databases and data access middleware and high-volume batch and (near) real-time processing