A prominent tech company is seeking an experienced Data Engineer to develop and manage robust ETL pipelines using Apache Spark (Scala). Candidates should possess 8 years of experience and be skilled in building scalable and reliable data solutions. Responsibilities include managing data workflows on AWS or Azure and optimizing data pipeline performance. The ideal candidate will have solid experience with Python, SQL, and big data technologies, along with excellent communication skills.
Qualifications
8 years of experience in Data Engineering.
Hands-on experience with data processing pipelines.
Ability to optimize data pipeline performance and quality.
Responsibilities
Develop and manage robust ETL pipelines using Apache Spark (Scala).
Collaborate cross-functionally to design effective data solutions.
Monitor, troubleshoot, and optimize pipeline performance.
Skills
Building solutions in Big Data environments
Data pipelines (scalable, fault tolerant, reliable)
Python
Apache Spark
Kafka
AWS services
Azure services
SQL
NoSQL technologies
Data Warehousing and ETL concepts
High-Level Design (HLD)
Low-Level Design (LLD)
Communication skills
Job description
Data Engineer
Experience: 8 years
Roles and Responsibilities
Develop and manage robust ETL pipelines using Apache Spark (Scala)
Understand Spark concepts, performance optimization techniques, and governance tools
Develop a highly scalable, reliable, and high-performance data processing pipeline to extract, transform, and load data from various systems to the Enterprise Data Warehouse/Data Lake/Data Mesh hosted on AWS or Azure
Collaborate cross-functionally to design effective data solutions
Implement data workflows utilizing AWS Step Functions or Azure Logic Apps for efficient orchestration. Leverage AWS Glue and Crawler or Azure Data Factory and Data Catalog for seamless data cataloging and automation
Monitor, troubleshoot, and optimize pipeline performance and data quality
Maintain high coding standards and produce thorough documentation. Contribute to high-level (HLD) and low-level (LLD) design discussions
Technical Skills
Progressive experience building solutions in Big Data environments
Have a strong ability to build robust and resilient data pipelines which are scalable, fault tolerant, and reliable in terms of data movement
Hands‑on expertise in Python, Spark, and Kafka
Strong command of AWS or Azure services
Strong hands‑on capabilities on SQL and NoSQL technologies
Sound understanding of data warehousing, modeling, and ETL concepts
Familiarity with High-Level Design (HLD) and Low-Level Design (LLD) principles