Data Engineer

Insight Global

Bengaluru Urban

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insight Global is seeking a skilled Data Engineer to design, build, and optimize large-scale batch data pipelines in a cloud environment. The role emphasizes reliability, performance, and data quality to support analytics and downstream consumers.

Candidates should have strong experience with Spark, Hadoop, Hive, and SQL across BigQuery and Spark SQL, plus hands-on use of GCP services like BigQuery, Dataproc, Pub/Sub, and GCS. Familiarity with Airflow or Cloud Composer is a plus.

Qualifications

  • Strong experience with Apache Spark, Hadoop, and Hive.
  • Hands-on experience building batch data pipelines with a focus on performance, scalability, SLA adherence, and fault tolerance.
  • Strong programming skills in Scala, with deep experience using Spark for data processing and analytics.
  • Experience working with GCP services including BigQuery, Google Cloud Storage (GCS), Dataproc, and Pub/Sub.
  • Solid experience writing and optimizing SQL, preferably BigQuery SQL and/or Spark SQL.
  • Strong understanding of data modeling, ETL/ELT patterns, and data quality best practices.
  • Experience with Kafka or similar messaging/streaming platforms.
  • Familiarity with workflow orchestration tools (e.g., Airflow or Cloud Composer).
  • Experience deploying and operating data pipelines in production cloud environments (GCP preferred, Azure acceptable).
  • Strong troubleshooting skills and ability to optimize pipelines under real-world constraints.

Responsibilities

  • Design, develop, and maintain batch data pipelines in a cloud environment using Spark, Hadoop, Hive.
  • Build fault-tolerant, SLA-driven pipelines that scale reliably.
  • Use GCP services (BigQuery, GCS, Dataproc, Pub/Sub) for ingestion and storage.
  • Write and optimize SQL (BigQuery/Spark SQL) for analysis and performance.
  • Collaborate with analytics, data science, and downstream teams to ensure data availability.
  • Monitor pipelines, troubleshoot failures, and implement alerting and retries.
  • Improve performance via partitioning, clustering, tuning, and query optimization.
  • Follow software practices: version control, testing, and documentation.

Skills

Apache Spark
Hadoop
Hive
Scala
Spark SQL
SQL
Data modeling
ETL/ELT
Data quality
Troubleshooting
Kafka
Airflow
Cloud Composer
GCS
BigQuery
Dataproc
Pub/Sub
Production deployments

Tools

BigQuery
GCS
Dataproc
Pub/Sub
Kafka
Airflow
Cloud Composer

Job description

Strong experience with big data technologies such as Apache Spark, Hadoop, and Hive.

Hands-on experience building batch data pipelines with a focus on performance, scalability, SLA adherence, and fault tolerance.

Strong programming skills in Scala, with deep experience using Spark for data processing and analytics.

Experience working with GCP services including BigQuery, Google Cloud Storage (GCS), Dataproc, and Pub/Sub.

Solid experience writing and optimizing SQL, preferably BigQuery SQL and/or Spark SQL.

Strong understanding of data modeling, ETL/ELT patterns, and data quality best practices.

Experience with Kafka or similar messaging/streaming platforms.

Familiarity with workflow orchestration tools (e.g., Airflow or Cloud Composer).

Experience deploying and operating data pipelines in production cloud environments (GCP preferred, Azure acceptable).

Strong troubleshooting skills and ability to optimize pipelines under real-world constraints.

Job Description

We are seeking a skilled Data Engineer to design, build, and optimize large-scale batch data pipelines in a cloud environment. This role focuses on reliability, performance, and data quality, supporting analytics and downstream consumers through well-engineered big-data solutions. The ideal candidate has strong experience with Apache Spark, cloud data platforms (GCP preferred), and writing performant SQL at scale.

Key Responsibilities
  • Design, develop, and maintain batch data pipelines using Apache Spark, Hadoop, Hive, or similar frameworks in a cloud environment.
  • Build highly optimized, fault-tolerant, and SLA-driven data pipelines that operate reliably at scale.
  • Leverage Google Cloud Platform (GCP) services such as BigQuery, GCS, Dataproc, and Pub/Sub to support data ingestion, processing, and storage.
  • Write and optimize SQL queries (BigQuery SQL and/or Spark SQL) for data analysis, profiling, and performance tuning.
  • Collaborate closely with analytics, data science, and downstream consumers to ensure data availability, correctness, and usability.
  • Monitor and troubleshoot pipeline failures; implement alerting, retries, and data quality checks.
  • Improve pipeline performance through partitioning, clustering, resource tuning, and query optimization.
  • Follow software engineering best practices, including version control, testing, and documentation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Gcp Data Engineer
Gcp Data Engineer

Pyramid It Consulting • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Insight Global • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Data Engineer
Data Engineer

EXL • India

On-site
INR 900,000 - 1,300,000
Data Engineer – GCP Java & Big Data
Data Engineer – GCP Java & Big Data

Impetus • Chennai District

On-site
INR 3,200,000 - 5,200,000
Big Data Developer
Big Data Developer

Objectways • Chennai District

On-site
INR 1,000,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Expian Technologies • Bengaluru

Hybrid
INR 1,000,000 - 1,500,000
Hiring For GCP Data Engineer For PAN India
Hiring For GCP Data Engineer For PAN India

HCLTech • Dadri, Chennai District, Bengaluru

On-site
INR 2,500,000 - 4,200,000
Sr. Data Engineer
Sr. Data Engineer

BigThinkCode • Chennai District

On-site
INR 800,000 - 1,200,000
Data Engineer II - GCP Data Platform
Data Engineer II - GCP Data Platform

Deutsche Telekom Digital Labs • Gurugram District

On-site
INR 1,500,000 - 2,100,000
Data Engineer-II
Data Engineer-II

Questhiring • Gurugram District

On-site
INR 1,500,000 - 2,800,000