Spark + Scala+ Python +Github Copilot

SPG Consulting

Bengaluru Urban

On-site

INR 1,200,000 - 2,300,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SPG Consulting in Bengaluru, India is seeking an experienced Data Engineer with 4–8 years of exp. You will design, develop and maintain scalable data pipelines using Apache Spark, Scala and PySpark, and optimize Spark workloads across distributed platforms.

You will leverage GitHub Copilot for AI-assisted development, build batch and near-real-time solutions, and collaborate with cross-functional teams in an Agile environment.

Qualifications

  • Strong hands-on experience with Apache Spark.
  • Strong programming experience in Scala.
  • Good hands-on experience with Python/PySpark.
  • Strong understanding of Spark SQL, DataFrames, and RDDs.
  • Experience with distributed computing and large-scale data processing.
  • Strong knowledge of data structures, algorithms, and performance optimization.
  • Experience with Git and GitHub.
  • Hands-on experience with GitHub Copilot or similar AI-assisted development tools.
  • Good understanding of ETL/ELT concepts and data engineering principles.
  • Strong debugging and problem-solving skills.
  • Good communication and collaboration skills.

Responsibilities

  • Design, develop, and maintain scalable data processing pipelines using Apache Spark.
  • Develop high-performance Spark applications using Scala and Python (PySpark).
  • Build batch and near-real-time data processing solutions.
  • Perform data transformation, cleansing, aggregation, and enrichment using Spark.
  • Optimize Spark jobs for performance, scalability, memory utilization, and cost.
  • Work with large datasets across distributed data platforms.
  • Develop reusable and maintainable Scala/Python code following coding standards.
  • Troubleshoot Spark jobs, performance issues, data-quality problems, and production failures.
  • Implement data validation, error handling, logging, and monitoring.
  • Work with cloud-based data platforms and distributed storage systems.
  • Use Git/GitHub for source control, branching, code reviews, and collaboration.
  • Leverage GitHub Copilot for code generation, refactoring, unit tests, documentation, SQL, and development productivity.
  • Review and validate Copilot-generated code for correctness, security, performance, and maintainability.
  • Collaborate with Data Architects, Data Engineers, Analysts, and business stakeholders.
  • Participate in Agile ceremonies, technical discussions, code reviews, and production support.

Skills

Apache Spark
Scala
Python
GitHub Copilot
Git
GitHub
Spark SQL
PySpark
DataFrames
HDFS

Education

Bachelor's or Master’s degree in Computer Science or related field

Tools

Azure Databricks
AWS EMR
Google Cloud Dataproc
Delta Lake
Kafka
Hadoop
Docker
Kubernetes
CI/CD

Job description

Bangalore North, India | Posted on 08/19/2026

Position
Experience

4–8 years

Job Summary

We are looking for an experienced Data Engineer with strong expertise in Apache Spark, Scala, Python, and GitHub Copilot. The candidate will be responsible for developing scalable data processing solutions, building data pipelines, optimizing Spark workloads, and leveraging AI-assisted development tools to improve engineering productivity and code quality.

Key Responsibilities

Design, develop, and maintain scalable data processing pipelines using Apache Spark.

Develop high-performance Spark applications using Scala and Python (PySpark).

Build batch and near-real-time data processing solutions.

Perform data transformation, cleansing, aggregation, and enrichment using Spark.

Optimize Spark jobs for performance, scalability, memory utilization, and cost.

Work with large datasets across distributed data platforms.

Develop reusable and maintainable Scala/Python code following coding standards.

Troubleshoot Spark jobs, performance issues, data-quality problems, and production failures.

Implement data validation, error handling, logging, and monitoring.

Work with cloud-based data platforms and distributed storage systems.

Use Git/GitHub for source control, branching, code reviews, and collaboration.

Leverage GitHub Copilot for code generation, refactoring, unit tests, documentation, SQL, and development productivity.

Review and validate Copilot-generated code for correctness, security, performance, and maintainability.

Collaborate with Data Architects, Data Engineers, Analysts, and business stakeholders.

Participate in Agile ceremonies, technical discussions, code reviews, and production support.

Required Skills

Strong hands-on experience with Apache Spark.

Strong programming experience in Scala.

Good hands-on experience with Python/PySpark.

Strong understanding of Spark SQL, DataFrames, and RDDs.

Experience with distributed computing and large-scale data processing.

Strong knowledge of data structures, algorithms, and performance optimization.

Experience with Git and GitHub.

Hands-on experience with GitHub Copilot or similar AI-assisted development tools.

Good understanding of ETL/ELT concepts and data engineering principles.

Strong debugging and problem-solving skills.

Good communication and collaboration skills.

Good to Have

Experience with Azure Databricks, AWS EMR, or Google Cloud Dataproc.

Knowledge of Delta Lake / Delta Tables.

Experience with Apache Kafka or other streaming technologies.

Knowledge of Hive, HDFS, and Hadoop ecosystem.

Experience with Azure Data Factory, AWS Glue, or similar orchestration tools.

Knowledge of Docker and Kubernetes.

Experience with CI/CD and DevOps practices.

Knowledge of SQL and relational databases.

Experience with cloud data warehouses such as Snowflake, Azure Synapse, or BigQuery.

Use GitHub Copilot to accelerate development of Scala and Python applications.

Generate and enhance unit tests, documentation, and repetitive code.

Use Copilot for debugging, refactoring, and code optimization.

Apply proper engineering judgment when reviewing AI-generated code.

Ensure generated code complies with organizational security, coding, and data-governance standards.

Education

Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field.

Key Technologies

Apache Spark | Scala | Python | PySpark | Spark SQL | Git | GitHub | GitHub Copilot | Databricks | Kafka | Hadoop | Cloud

Preferred Candidate Profile

The ideal candidate should have strong hands-on experience in Spark, Scala, and Python, with a solid understanding of distributed data processing and modern data engineering practices. Experience using GitHub Copilot effectively and responsibly to improve development productivity is highly desirable.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Spark + Scala+ Python + Github + Copilot
Spark + Scala+ Python + Github + Copilot

Hexaware Technologies • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer (Spark/Scala)
Data Engineer (Spark/Scala)

Zorba AI • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Contractor - PySpark Engineer
Contractor - PySpark Engineer

Vivantify • Hyderabad

On-site
INR 1,800,000 - 2,600,000
Databricks Architect
Databricks Architect

SPG Consulting • Bengaluru Urban

On-site
INR 4,500,000 - 6,500,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Data Engineer (APAC Region)
Senior Data Engineer (APAC Region)

ANRGI TECH • Maharashtra

On-site
INR 2,500,000 - 3,800,000
Opportunity for Data Engineer
Opportunity for Data Engineer

Hinduja Tech Limited • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior Data Engineer
Senior Data Engineer

Amgen • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Scala + Spark + SQL ( Data Engineer )
Scala + Spark + SQL ( Data Engineer )

3minds Esolutions • Bengaluru, Bangalore Rural, Hyderabad

On-site
INR 900,000 - 1,400,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000