Data Engineer GCP Java & Big Data

ClearTrail Technologies

Gurugram District

On-site

INR 1,200,000 - 2,400,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ClearTrail Technologies in Gurugram, India is seeking a skilled GCP Data Engineer with extensive hands-on experience in building scalable, high-performance data solutions using Java-based big data processing frameworks on Google Cloud Platform.

The role focuses on designing, developing, and maintaining ETL/ELT pipelines with Dataproc, Dataflow, Spark, BigQuery, and related GCP services, while ensuring robust, fault-tolerant data architectures and collaboration with cross-functional teams.

Qualifications

  • Strong knowledge of Java, data structures, and algorithms.
  • Experience with Spark, Hadoop, and Hive.
  • Proficient in SQL and data modeling for data warehousing.
  • Unix/Linux scripting experience.
  • Ability to design scalable, fault-tolerant data processing systems.
  • Experience with GCP services (BigQuery, Dataflow/Dataproc, Pub/Sub).

Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using Java-based big data frameworks on GCP
  • Build and optimize batch and streaming data processing solutions using Dataproc, Dataflow, and Spark
  • Ensure high-quality, efficient, and maintainable code by following best practices
  • Perform unit testing and troubleshoot data pipeline issues
  • Collaborate with cross-functional teams to understand data requirements and deliver robust solutions
  • Design cost-optimized and performance-efficient architectures leveraging GCP-native services

Skills

Java
Apache Spark
Hadoop
Hive
SQL / Data Modeling
Linux / Bash
Distributed Systems
GCP services

Job description

Job Summary:


We are looking for skilled GCP Data Engineers with 69 years of hands‑on experience in building scalable, high-performance data solutions. The ideal candidate will have strong expertise in Java-based big data processing frameworks and deep exposure to Google Cloud Platform (GCP) services for modern data engineering workloads.


Key Skills & Experience:


  1. Strong programming expertise in Java, with experience in building distributed data processing applications

  2. Hands‑on experience with Big Data technologies such as Apache Spark (Java APIs), Hadoop, and Hive

  3. Experience with Spark (DataFrame/Spark SQL) using Java (PySpark knowledge

  4. Solid understanding of data structures, algorithms, and object‑oriented programming in Java

  5. Strong knowledge of SQL, data modeling, and data warehousing concepts

  6. Experience working with Linux/Unix environments and scripting (Bash or similar)

  7. Proven analytical and problem‑solving skills, especially in debugging and optimizing data pipelines

  8. Ability to design and build scalable, fault‑tolerant data processing systems


Preferred / Good to Have:


  1. Hands‑on experience with GCP services such as BigQuery, Dataflow (Apache Beam with Java), Dataproc, Cloud Storage, Pub/Sub, and IAM

  2. Experience with workflow orchestration tools like Airflow or Cloud Composer

  3. Exposure to cloud migration projects, especially transitioning from on‑premise Hadoop ecosystems to GCP

  4. Familiarity with streaming data pipelines using Pub/Sub and Dataflow

  5. Understanding of CI/CD pipelines and DevOps practices in a cloud environment


Roles & Responsibilities:


  1. Design, develop, and maintain scalable ETL/ELT pipelines using Java‑based big data frameworks on GCP

  2. Build and optimize batch and streaming data processing solutions using Dataproc, Dataflow, and Spark

  3. Ensure high‑quality, efficient, and maintainable code by following best practices and coding standards

  4. Perform unit testing, integration testing, and troubleshoot complex data pipeline issues

  5. Collaborate with cross‑functional teams to understand data requirements and deliver robust solutions

  6. Estimate development efforts and contribute to sprint planning and delivery

  7. Participate in code reviews and mentor junior team members where required

  8. Design cost‑optimized and performance‑efficient architectures leveraging GCP‑native services


Roles and Responsibilities


Skills : Java, Bigdata ,GCP
Core Java: Well versed with OOP, Data Structures, Generics, Collections, Basic Regular Expressions, IO, Basic Concurrency. Java Spring: Core, REST API's Knowledge of: Basics Shell scripting, Postman, JSON, MYSQL Big Data: Hadoop, Map reduce, Basic Spark, HBase(M7) GCP skillset, Big query


We are looking for a skilled Software Engineer / Data Engineer with strong expertise in Core Java, Big Data technologies, and GCP to design, develop, and maintain scalable data processing systems and microservices.


Primary Skills / Technical ExpertiseCore Java



  • Strong knowledge of OOP concepts, Data Structures & Algorithms

  • Expertise in Collections, Generics, Exception Handling

  • Experience in Multithreading & Basic Concurrency

  • Hands‑on with Java IO & file processing

  • Understanding of Regular Expressions


Java & Spring Framework



  • Experience in Spring Core, Spring Boot

  • Strong exposure to REST API design and development

  • Knowledge of Microservices architecture


Big Data Technologies



  • Hands‑on experience with:

    • Hadoop Ecosystem

    • MapReduce

    • Apache Spark (Core & Basics of Spark SQL/PySpark)

    • Hive (good to have)

    • HBase or other NoSQL systems


  • Understanding of:

    • Distributed computing concepts

    • Batch & real-time data processing pipelines


  • Ability to handle large-scale data (GBs to TBs) (typical in Big Data roles) [Round 2 MI...Recording | Video]


GCP (Google Cloud Platform)



  • Hands‑on experience with:

    • BigQuery

    • Cloud Storage (GCS)


  • Good to have:

    • Dataflow / Dataproc

    • Pub/Sub


  • Understanding of cloud-based data pipelines and deployment


Database / Tools



  • Strong knowledge of:

    • SQL (MySQL / Hive / BigQuery queries)

    • JSON handling & API integration


  • Tools:

    • Postman (API testing)

    • Shell scripting (basic automation)

    • Version control tools (Git – optional but preferred)



Key Responsibilities



  • Design and develop scalable data processing applications

  • Build and optimize ETL/data pipelines for large datasets

  • Develop and maintain REST APIs and microservices

  • Work on Big Data transformations and storage solutions

  • Integrate systems with GCP cloud services

  • Perform data validation, testing, and performance tuning

  • Collaborate with cross-functional teams for end‑to‑end delivery

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer – GCP Java & Big Data
Data Engineer – GCP Java & Big Data

Impetus • Chennai District

On-site
INR 3,200,000 - 5,200,000
Data Engineering-ETL-Big Data-GCP-Java
Data Engineering-ETL-Big Data-GCP-Java

BCforward • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Data Engineer – GCP Java & Big Data
Data Engineer – GCP Java & Big Data

United States Digital Space LLC • Karnataka

On-site
INR 2,500,000 - 4,500,000
Hiring For GCP Data Engineer For PAN India
Hiring For GCP Data Engineer For PAN India

HCLTech • Dadri, Chennai District, Bengaluru

On-site
INR 2,500,000 - 4,200,000
Gcp Data Engineer
Gcp Data Engineer

Lloyds Technology Centre • Hyderabad

Hybrid
INR 2,500,000 - 4,200,000
GCP Data Engineer
GCP Data Engineer

EXL • Gurugram District

On-site
INR 1,500,000 - 2,100,000
GCP Data Engineer
GCP Data Engineer

Elabs Infotech • Bengaluru

Hybrid
INR 1,200,000 - 2,400,000
Data Engineer
Data Engineer

EXL • India

On-site
INR 900,000 - 1,300,000
GCP Data Engineer
GCP Data Engineer

Religent Systems • Hyderabad

On-site
INR 1,200,000 - 1,800,000
DATA ENGINEER
DATA ENGINEER

Covaicareers.com • Chennai District

Hybrid
INR 2,500,000 - 4,000,000