Data Engineer – GCP Java & Big Data

Impetus

Chennai District

On-site

INR 3,200,000 - 5,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Impetus in Chennai, India, seeks a skilled GCP Data Engineer with 6–9 years of hands-on experience to design scalable data pipelines on Google Cloud Platform using Java-based big‑data processing frameworks.

You will build batch and streaming data solutions with Dataproc, Dataflow, Spark and BigQuery, ensure robust, cost-efficient architectures, mentor juniors and collaborate with cross‑functional teams to deliver data products.

Qualifications

  • Experience in Java-based big data frameworks on GCP required.
  • Strong SQL, data modeling and warehousing knowledge.
  • Hands-on with Spark, Hadoop, Hive and data pipelines.

Responsibilities

  • Design, develop and maintain scalable ETL/ELT pipelines in Java-based big data frameworks on GCP.
  • Build batch and streaming data solutions using Dataproc, Dataflow and Spark.
  • Ensure high-quality, maintainable code and participate in code reviews.

Skills

Java core
Big data (Spark)
SQL & data warehousing
Linux scripting
Data pipelines
GCP (BigQuery, Dataflow)
Airflow/Cloud Composer
CI/CD/DevOps
Problem solving

Tools

Hadoop
Hive
Spark
Airflow
Dataproc
BigQuery

Job description

Job Summary

We are looking for skilled GCP Data Engineers with 6–9 years of hands‑on experience building scalable, high‑performance data solutions. The ideal candidate will have strong expertise in Java‑based big‑data processing frameworks and deep exposure to Google Cloud Platform (GCP) services for modern data engineering workloads.

Key Skills & Experience
  • Strong programming expertise in Java, building distributed data processing applications
  • Hands‑on experience with Big Data technologies such as Apache Spark (Java/Scala APIs), Hadoop and Hive
  • Experience with Spark DataFrames/Spark SQL using Java or Scala (PySpark knowledge is a plus)
  • Solid understanding of data structures, algorithms and OOP in Java
  • Strong knowledge of SQL, data modelling and data warehousing concepts
  • Experience working with Linux/Unix environments and scripting (Bash or similar)
  • Proven analytical and problem‑solving skills, especially in debugging and optimizing data pipelines
  • Ability to design and build scalable, fault‑tolerant data processing systems
Important To Have
  • Hands‑on experience with GCP services such as BigQuery, Dataflow (Apache Beam with Java), Dataproc, Cloud Storage, Pub/Sub and IAM
  • Experience with workflow orchestration tools like Airflow or Cloud Composer
  • Exposure to cloud migration projects, particularly transitioning from on‑premise Hadoop ecosystems to GCP
  • Familiarity with streaming data pipelines using Pub/Sub and Dataflow
  • Understanding of CI/CD pipelines and DevOps practices in a cloud environment
Roles & Responsibilities
  • Design, develop and maintain scalable ETL/ELT pipelines using Java‑based big data frameworks on GCP
  • Build and optimise batch and streaming data processing solutions using Dataproc, Dataflow and Spark
  • Ensure high‑quality, efficient and maintainable code by following best practices and coding standards
  • Perform unit and integration testing and troubleshoot complex data pipeline issues
  • Collaborate with cross‑functional teams to understand data requirements and deliver robust solutions
  • Estimate development efforts and contribute to sprint planning and delivery
  • Participate in code reviews and mentor junior team members where required
  • Design cost‑optimised and performance‑efficient architectures leveraging GCP‑native services
Technical Expertise
  • Core Java
    • OOP concepts, data structures & algorithms
    • Collections, generics, exception handling
    • Multithreading & basic concurrency
    • Java I/O & file processing
    • Regular expressions
  • Java & Spring Framework
    • Spring Core, Spring Boot
    • REST API design and development
    • Microservices architecture
  • Big Data Technologies
    • Hadoop ecosystem, MapReduce
    • Apache Spark (core, Spark SQL, PySpark)
    • Hive (good to have)
    • HBase or other NoSQL systems
    • Distributed computing concepts, batch & real‑time pipelines
  • Google Cloud Platform
    • BigQuery, Cloud Storage (GCS)
    • Dataflow, Dataproc, Pub/Sub
    • Understanding of cloud‑based data pipelines and deployment
  • Database / Tools
    • SQL (MySQL, Hive, BigQuery)
    • JSON handling & API integration
    • Postman (API testing)
    • Shell scripting (basic automation)
    • Version control (Git – optional but preferred)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer – GCP Java & Big Data
Data Engineer – GCP Java & Big Data

United States Digital Space LLC • Karnataka

On-site
INR 2,500,000 - 4,500,000
Data Engineer GCP Java & Big Data
Data Engineer GCP Java & Big Data

ClearTrail Technologies • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Data Engineering-ETL-Big Data-GCP-Java
Data Engineering-ETL-Big Data-GCP-Java

BCforward • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Hiring For GCP Data Engineer For PAN India
Hiring For GCP Data Engineer For PAN India

HCLTech • Dadri, Chennai District, Bengaluru

On-site
INR 2,500,000 - 4,200,000
Gcp Data Engineer
Gcp Data Engineer

Synoptek • Pune District, Ahmedabad District

Hybrid
INR 900,000 - 1,500,000
GCP Data Engineer
GCP Data Engineer

Tata Consultancy Services • Tamil Nadu

On-site
INR 1,500,000 - 3,200,000
Gcp Data Engineer
Gcp Data Engineer

Pyramid It Consulting • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer - GCP (Google Cloud Platform)
Data Engineer - GCP (Google Cloud Platform)

HCLTech • Hyderabad

Hybrid
INR 2,400,000 - 4,200,000
GCP Data Engineer_Pyspark/Scala
GCP Data Engineer_Pyspark/Scala

Zorba AI • Chennai District

On-site
INR 400,000 - 640,000
GCP Data Engineer
GCP Data Engineer

Religent Systems • Hyderabad

On-site
INR 1,200,000 - 1,800,000