A leading data solutions firm is seeking a Data Engineer to design and implement Big Data solutions using Hadoop, AWS, and other cloud technologies. The ideal candidate should have over 6 years of data engineering experience, proficiency in programming languages like Java, Scala, or Python, and a solid understanding of SQL and NoSQL databases. This role requires familiarity with various data handling frameworks and the ability to optimize data pipelines, ensuring scalability and efficiency in data processing.
Qualifications
6+ years of Data engineering experience.
Good hands-on knowledge of Java, Scala, or Python.
Understanding of SQL & NoSQL databases like MySQL and Mongo.
Familiar with cloud solutions, preferably AWS.
Hands-on experience with data handling frameworks like Spark.
Responsibilities
Design data solutions using Hadoop based technologies.
Implement design and various components of Big Data platforms.
Create and optimize data access layers for SQL queries.
Develop scalable solutions using big data/cloud technologies.
Skills
Java
Scala
Python
SQL
NoSQL
Spark
Kafka
AWS
Tools
Hadoop
Apache Beam
Apache Flink
Job description
Responsibilities
You will be involved in the design of data solutions using Hadoop based technologies along with Hadoop, AWS.
Responsibilities includes design and implementation of various Big Data platform components like (Batch Processing, Live Stream Processing, In----Memory Cache, Query Layer (SQL), Rule Engine and Action Framework ).
Design and Implemented Data Access Layer, which can connect to various data sources and uses advanced caching techniques to provide fast responses to real time SQL queries using Big Data Technologies.
Implement scalable solutions to meet the ever-increasing data volumes, using big data/cloud technologies; spark, Kafka, any Cloud computing etc.
Requirements
6+ years of Data engineering experience.
Programming Language: Good hands-on knowledge on one of the programming language java/scala/python.
Data Store: Good understanding of SQL & NoSQL Data bases. Mysql, Mongo.
Data Infrastructure: Familiar with cloud solutions for data infrastructure. AWS Preferably.
Data handling frameworks: Good hands on one of the data handling frameworks - Spark, Apache Beam, Apache flink etc
File formats: Familiar different types of data formats - Apache Parquet, Avro, ORC etc.
Good understanding of how to setup and optimise data pipeline and overall infrastructure for setting up ETL jobs.
Proficient with distributed file system and computations concepts.