Spark + Scala+ Python + Github + Copilot

Hexaware Technologies

Bengaluru

Hybrid

INR 1,200,000 - 1,800,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hexaware Technologies in Bengaluru is seeking a Senior Data Engineer with 7+ years of experience to design and maintain large-scale data pipelines using Spark, PySpark, Scala, Python, Hive and Airflow. The ideal candidate should have hands-on experience with HDFS, Parquet, and data modeling concepts such as Star and Snowflake schemas.

Experience with Cloudera CDP on-prem, CI/CD pipelines, and using modern AI-assisted tools to boost development productivity is highly preferred.

Qualifications

  • 7+ years of experience developing and maintaining large-scale data processing solutions using Spark, PySpark, Scala, Python, Hive, and Airflow.
  • N hands-on with Cloudera CDP (On-Premise) environments is highly preferred.
  • Experience with Hive SQL, partitioning, bucketing, performance tuning, and ETL workflows.
  • Strong knowledge of Spark SQL, DataFrames, Datasets, UDFs, and Spark optimization techniques.
  • Exposure to HDFS, Parquet, data modeling concepts such as Star and Snowflake schemas; familiarity with Git, GitLab, Maven/SBT, CI/CD pipelines.

Responsibilities

  • Design, develop, and maintain enterprise-scale Spark/PySpark data pipelines using Scala and Python.
  • Work with HDFS, Parquet, and large-scale distributed data environments.
  • Optimize Spark, Hive, and ETL workloads for performance, scalability, and reliability.
  • Implement CI/CD pipelines and DevOps best practices using GitLab, Jenkins, Maven, and SBT.
  • Utilise AI-powered development tools to improve code quality, engineering productivity, and development efficiency.

Skills

ETL development
Distributed computing
Data modeling
Workflow orchestration
Automation

Tools

Spark
PySpark
Scala
Python
Hive
Apache Airflow
HDFS
Parquet
Cloudera CDP
Git
GitLab
Maven/SBT
CI/CD
Impala
Kafka
Power BI
GitHub Copilot
Cursor
Claude
Hive SQL
Star/Snowflake schemas

Job description

Senior Data Engineer - Spark / PySpark / Scala/ Hive Experience: 7+ Years Location: Bangalore (Hybrid)
Job Summary

We are looking for a highly skilled Senior Data Engineer with 7+ years of experience in developing and maintaining large-scale data processing solutions using Spark, PySpark, Scala, Python, Hive, and Apache Airflow. The ideal candidate should have strong expertise in distributed computing, ETL development, DataLake, workflow orchestration, and automation. Hands-on experience with Cloudera CDP (On-Premise) environments is highly preferred. The candidate should also be comfortable using modern AI-assisted development tools such as Claude, Cursor, GitHub Copilot to improve development productivity and code quality. We are looking for a highly skilled Big Data Engineer with expertise in Apache Spark, PySpark, Scala, Python, Hive, and Apache Airflow to build and support enterprise-scale data platforms. The role requires hands-on experience with Hive SQL, partitioning, bucketing, performance tuning, and ETL workflows, along with strong knowledge of Spark SQL, DataFrames, Datasets, UDFs, and Spark optimization techniques. Experience with HDFS, Parquet, Cloudera CDP, and data modeling concepts such as Star and Snowflake schemas is essential. Candidates should have experience with Git, GitLab, Maven/SBT, CI/CD pipelines. Experience with GitHub Copilot, Cursor, or Claude, as well as exposure to Impala, Kafka, and Power BI, is a plus.

Key Responsibilities
  • Design, develop, and maintain enterprise-scale Spark/PySpark data pipelines using Scala and Python.
  • Work with HDFS, Parquet, and large-scale distributed data environments.
  • Optimize Spark, Hive, and ETL workloads for performance, scalability, and reliability.
  • Implement CI/CD pipelines and DevOps best practices using GitLab, Jenkins, Maven, and SBT.
  • Utilise AI-powered development tools to improve code quality, engineering productivity, and development efficiency.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Spark + Scala+ Python +Github Copilot
Spark + Scala+ Python +Github Copilot

SPG Consulting • Bengaluru Urban

On-site
INR 1,200,000 - 2,300,000
Data Engineer (PySpark + Cloudera)
Data Engineer (PySpark + Cloudera)

Zorba AI • Maharashtra

On-site
INR 1,000,000 - 1,500,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Senior Data Engineer
Senior Data Engineer

Arcana Analytics • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer
Senior Data Engineer

N Consulting Limited • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Big Data Developer
Big Data Developer

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Senior Data Engineer
Senior Data Engineer

Moolya Software Testing • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Data Engineer (Spark/Scala)
Data Engineer (Spark/Scala)

Zorba AI • Hyderabad

On-site
INR 1,800,000 - 3,000,000