A leading technology firm is seeking an experienced Data Engineer to create and maintain data pipeline architecture. The ideal candidate will have over 8 years of experience in data engineering, particularly in machine learning projects, and proficiency with Azure cloud services and big data tools like Hadoop and Spark. This role is essential for enhancing data delivery and supporting large data transformation initiatives, contributing to the overall success of the project.
Qualifications
8+ years of experience as a data engineer focusing on machine learning.
Proficient in building and optimizing data pipelines and datasets.
Experience with big data tools and ETL processes.
Responsibilities
Create and maintain optimal data pipeline architecture.
Design production data pipelines from ingestion to consumption.
Implement internal process improvements for data handling.
Skills
Data pipeline architecture
Analytic skills
Team collaboration
Education
Bachelor's or master's degree in computer science or related field
Tools
Azure cloud services
Spark
Hadoop
Airflow
Job description
Responsibilities
Create and maintain optimal data pipeline architecture; assemble large, complex data sets that meet functional/non‑functional requirements.
Design the right schema to support the functional requirement and consumption pattern.
Design and build production data pipelines from ingestion to consumption.
Create necessary preprocessing and postprocessing for various forms of data for training/retraining and inference ingestions as required.
Create data visualization and business intelligence tools for stakeholders and data scientists for necessary business/solution insights.
Identify, design, and implement internal process improvements: automating manual data processes, optimizing data delivery, etc.
Ensure our data is separated and secure across national boundaries through multiple data centers.
Requirements & Skills
You should have a bachelor’s or master’s degree in computer science, information technology or other quantitative fields.
You should have at least 8 years of experience as a data engineer supporting large data transformation initiatives related to machine learning, with experience in building and optimizing pipelines and data sets.
Strong analytic skills related to working with unstructured datasets.
Experience with Azure cloud services, ADF, ADLS, HDInsight, Data Bricks, App Insights, etc.
Experience handling ETL’s using Spark.
Experience with object‑oriented/ function scripting languages: Python, Pyspark, etc.
Experience with big data tools: Hadoop, Spark, Kafka, etc.
Experience with data pipeline and workflow management tools: Azkaban, Luigi, Airflow, etc.
You should be a good team player and committed to the success of the team and overall project.