Data Engineer - INTL India

Insight Global

Bentonville (AR)

On-site

USD 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insight Global is seeking a skilled Data Engineer to design, build, and optimize large-scale batch data pipelines in a cloud environment, with emphasis on reliability and data quality. The role focuses on Spark, Hadoop/Hive, Python/Scala, and GCP services like BigQuery, GCS, Dataproc, and Pub/Sub to support analytics and downstream consumers.

This on-site role requires collaboration with analytics teams, monitoring, and applying best practices in version control, testing, and documentation;

Qualifications

  • Strong experience with big data technologies such as Apache Spark, Hadoop, and Hive.
  • Hands-on experience building batch data pipelines with a focus on performance, scalability, SLA adherence, and fault tolerance.
  • Strong programming skills in Python and/or Scala, with deep experience using Spark for data processing and analytics.
  • Experience working with GCP services including BigQuery, Google Cloud Storage (GCS), Dataproc, and Pub/Sub.
  • Solid experience writing and optimizing SQL, preferably BigQuery SQL and/or Spark SQL.
  • Strong understanding of data modeling, ETL/ELT patterns, and data quality best practices.
  • Experience with Kafka or similar messaging/streaming platforms.
  • Familiarity with workflow orchestration tools (e.g., Airflow or Cloud Composer).
  • Experience deploying and operating data pipelines in production cloud environments (GCP preferred, Azure acceptable).
  • Strong troubleshooting skills and ability to optimize pipelines under real-world constraints.

Responsibilities

  • Design, develop, and maintain batch data pipelines using Apache Spark, Hadoop, Hive, or similar frameworks in a cloud environment.
  • Build highly optimized, fault-tolerant, and SLA-driven data pipelines that operate reliably at scale.
  • Leverage Google Cloud Platform (GCP) services such as BigQuery, GCS, Dataproc, and Pub/Sub to support data ingestion, processing, and storage.
  • Write and optimize SQL queries (BigQuery SQL and/or Spark SQL) for data analysis, profiling, and performance tuning.
  • Collaborate closely with analytics, data science, and downstream consumers to ensure data availability, correctness, and usability.
  • Monitor and troubleshoot pipeline failures; implement alerting, retries, and data quality checks.
  • Improve pipeline performance through partitioning, clustering, resource tuning, and query optimization.
  • Follow software engineering best practices, including version control, testing, and documentation.

Job description

Job Description

Must report onsite to the Bangalore Office 5 days a week**

We are seeking a skilled Data Engineer to design, build, and optimize large‑scale batch data pipelines in a cloud environment. This role focuses on reliability, performance, and data quality, supporting analytics and downstream consumers through well‑engineered big‑data solutions. The ideal candidate has strong experience with Apache Spark, cloud data platforms (GCP preferred), and writing performant SQL at scale.

Key Responsibilities
  • Design, develop, and maintain batch data pipelines using Apache Spark, Hadoop, Hive, or similar frameworks in a cloud environment.
  • Build highly optimized, fault‑tolerant, and SLA‑driven data pipelines that operate reliably at scale.
  • Leverage Google Cloud Platform (GCP) services such as BigQuery, GCS, Dataproc, and Pub/Sub to support data ingestion, processing, and storage.
  • Write and optimize SQL queries (BigQuery SQL and/or Spark SQL) for data analysis, profiling, and performance tuning.
  • Collaborate closely with analytics, data science, and downstream consumers to ensure data availability, correctness, and usability.
  • Monitor and troubleshoot pipeline failures; implement alerting, retries, and data quality checks.
  • Improve pipeline performance through partitioning, clustering, resource tuning, and query optimization.
  • Follow software engineering best practices, including version control, testing, and documentation.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Skills and Requirements
  • Strong experience with big data technologies such as Apache Spark, Hadoop, and Hive.
  • Hands‑on experience building batch data pipelines with a focus on performance, scalability, SLA adherence, and fault tolerance.
  • Strong programming skills in Python and/or Scala, with deep experience using Spark for data processing and analytics.
  • Experience working with GCP services including BigQuery, Google Cloud Storage (GCS), Dataproc, and Pub/Sub.
  • Solid experience writing and optimizing SQL, preferably BigQuery SQL and/or Spark SQL.
  • Strong understanding of data modeling, ETL/ELT patterns, and data quality best practices.
  • Experience with Kafka or similar messaging/streaming platforms.
  • Familiarity with workflow orchestration tools (e.g., Airflow or Cloud Composer).
  • Experience deploying and operating data pipelines in production cloud environments (GCP preferred, Azure acceptable).
  • Strong troubleshooting skills and ability to optimize pipelines under real‑world constraints.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Spark & GCP Pipelines at Scale (Bangalore)
Data Engineer - Spark & GCP Pipelines at Scale (Bangalore)

Insight Global • Bentonville (AR)

On-site
USD 110,000 - 160,000
Senior Data Engineer INDIA
Senior Data Engineer INDIA

Vytwo • Prosper (TX)

Remote
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 140,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Bentonville (AR)

On-site
USD 100,000 - 130,000
Data Engineer (in person)
Data Engineer (in person)

SEP • Westfield (IN)

On-site
USD 90,000 - 110,000
Flexible work schedules
Opportunities to learn and develop
Community of friendly peers
+1
Data Engineer
Data Engineer

TechDigital Group • Burbank (CA)

On-site
USD 120,000 - 150,000
Data Engineer
Data Engineer

Siro Clinpharm • United States

Remote
USD 37,000 - 73,000
Data Engineer
Data Engineer

VTG Defense • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

SAN R&D Business Solutions • Alpharetta (GA)

On-site
USD 110,000 - 140,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000