We are hiring a Big Data Engineer on behalf of our client. This role involves designing, building, and maintaining scalable data pipelines and infrastructure to process large volumes of structured and unstructured data. It's a great opportunity for someone who enjoys solving complex data engineering challenges and building systems that power data-driven decision-making across the organization.
Role & responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines for large-scale data processing
- Develop and optimize data lakes, data warehouses, and distributed systems
- Work with big data frameworks (Hadoop, Spark, Hive) to process massive datasets
- Build and manage real-time data streaming pipelines (Kafka, Flink, Spark Streaming)
- Ensure data quality, integrity, and governance across all pipelines
- Optimize data storage and retrieval for performance and cost efficiency
- Collaborate with data scientists, analysts, and engineering teams to meet data requirements
- Manage and monitor cloud-based big data infrastructure (AWS EMR, GCP Dataproc, Azure Synapse)
- Automate workflows using orchestration tools (Airflow, Luigi, Oozie)
- Troubleshoot and resolve data pipeline issues and performance bottlenecks
Preferred candidate profile
- 0 to 2 years of experience in big data engineering or related role
- Bachelor's/Master's degree in Computer Science, Engineering, or related field
- Strong programming skills in Python, Java, or Scala
- Hands-on experience with Hadoop ecosystem (HDFS, MapReduce, Hive, Pig)
- Strong expertise in Apache Spark (batch and streaming)
- Experience with SQL/NoSQL databases (MySQL, MongoDB, Cassandra, HBase)
- Familiarity with data streaming tools (Kafka, Flink)
- Experience with workflow orchestration (Airflow, Oozie, Luigi)
- Knowledge of cloud platforms (AWS, GCP, Azure) and their big data services
- Strong understanding of data modeling, data warehousing, and distributed computing concepts