Spark Expert/ Spark Developer

Datagaps

Hyderabad

Hybrid

INR 1,200,000 - 2,400,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Datagaps is looking for an Apache Spark Developer / Spark Performance Engineer to design, develop, optimize, and maintain large-scale data processing applications. You will focus on memory optimization, ETL pipelines, and distributed processing across data lakes and cloud storage.

The role requires 4+ years hands-on Spark development, strong knowledge of Spark architecture, and experience with Spark SQL/Streaming, Kubernetes, and cloud platforms.

Qualifications

  • Bachelor's degree in CS/IT/Engineering or related field.
  • 4+ years hands-on Apache Spark development experience.
  • Experience building enterprise-grade Big Data and Data Engineering solutions.

Responsibilities

  • Design, develop, and maintain scalable Spark data processing pipelines.
  • Optimize batch and real-time data processing with Spark SQL/Streaming.
  • Improve performance via memory management, query optimization, and resource tuning.
  • Troubleshoot bottlenecks, data skew, and executor failures.
  • Build robust ETL/ELT workflows across data lakes, Hadoop, and cloud storage.
  • Collaborate with Data Engineers, Architects, and business teams to design data solutions.
  • Monitor Spark apps using Spark UI, History Server, and cluster tools.
  • Apply partitioning, caching, and resource allocation best practices.
  • Tune Spark workloads on Kubernetes and cloud environments.
  • Support deployments with code reviews, testing, and documentation.

Skills

Spark architecture
RDD API
DataFrame API
Dataset API
Spark SQL
Spark Streaming
Kubernetes
Memory optimization
ETL pipelines
GC tuning
Data Lakes
Delta Lake
Hadoop
Hive
Airflow

Education

Bachelor's degree in CS/IT/Engineering

Tools

Spark UI
Spark History Server
Cluster Monitoring Tools
Databricks
Delta Lake
Hadoop
Hive
Kafka

Job description

Role Summary

We are seeking a highly skilled Apache Spark Developer / Spark Performance Engineer with 4+ years of hands‑on Spark experience to design, develop, optimize, and maintain large-scale data processing applications. The ideal candidate should possess strong expertise in Spark performance tuning, memory optimization, ETL pipelines, and distributed data processing while working across modern data lake and big data platforms.

Key Responsibilities
  • Design, develop, and maintain scalable data processing pipelines using Apache Spark.
  • Build and optimize batch and real‑time data processing solutions utilizing Spark SQL and Spark Streaming.
  • Analyze and improve Spark job performance through effective memory management, query optimization, and resource utilization.
  • Troubleshoot performance bottlenecks, data skew issues, shuffle inefficiencies, and executor failures.
  • Develop robust ETL/ELT workflows integrating data from multiple enterprise systems.
  • Work with large‑scale datasets stored in Data Lakes, Hadoop, HDFS, and cloud‑based storage platforms.
  • Collaborate with Data Engineers, Architects, and Business Teams to understand requirements and design optimal data solutions.
  • Monitor, maintain, and improve Spark applications using Spark UI, Spark History Server, and cluster monitoring tools.
  • Implement best practices for partitioning, caching, persistence, and resource allocation.
  • Optimize Spark workloads running on Kubernetes and cloud‑based environments.
  • Ensure code quality through reviews, testing, documentation, and adherence to development standards.
  • Support production deployments and provide performance tuning recommendations for existing Spark workloads.
Required Skills & Experience
Experience
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • Minimum 4+ years of hands‑on experience in Apache Spark Development.
  • Experience developing enterprise‑grade Big Data and Data Engineering solutions.
Core Apache Spark Expertise
  • Strong understanding of Spark Architecture and Distributed Computing concepts.
  • Hands‑on experience with:
    • RDD
    • DataFrame API
    • Dataset API
    • Spark SQL
    • Spark Streaming / Structured Streaming
    • Spark Session & Spark Context
    • DAG Scheduler
  • Deep understanding of:
    • Partitioning Strategies
    • Repartition and Coalesce
    • Broadcast Join
    • Shuffle Join
    • Query Optimization
Spark Performance Tuning & Monitoring (Mandatory)
  • Strong expertise in:
    • Executor Memory Tuning
    • Driver Memory Optimization
    • Garbage Collection (GC) Tuning
    • Dynamic Resource Allocation
    • Shuffle Optimization
    • Spill‑to‑Disk Reduction Techniques
  • Hands‑on experience using:
    • Spark UI
    • Spark History Server
    • Cluster Monitoring Tools
  • Experience handling:
    • Data Skew
    • Executor Failures
    • Performance Bottlenecks
    • Memory Leaks
  • Expertise in:
    • Caching and Persistence Strategies
    • MEMORY_ONLY
    • MEMORY_AND_DISK
    • Storage Optimization
Data Engineering & Storage
  • Strong experience building ETL and ELT pipelines.
  • Hands‑on experience with:
    • Data Lakes
    • Delta Lake
    • Hadoop Ecosystem
    • HDFS
    • Hive
  • Experience working with file formats:
    • Parquet
    • ORC
    • Avro
  • Understanding of Data Warehousing concepts and large‑scale data processing.
Cloud & Containerization
  • Experience running Spark workloads on Kubernetes.
  • Understanding of Spark‑on‑Kubernetes architecture and performance tuning.
  • Exposure to cloud platforms such as AWS, Azure, or GCP is preferred.
Nice‑to‑Have Skills
Databricks
  • Hands‑on experience with:
    • Databricks Platform
    • Delta Lake
    • Unity Catalog
    • Databricks Workflows
    • Databricks Jobs
    • Databricks Notebooks
  • Experience optimizing Spark workloads within Databricks environments.
Additional Skills
  • Knowledge of Airflow or workflow orchestration tools.
  • Experience with CI/CD pipelines and DevOps practices.
  • Familiarity with Kafka and real‑time data ingestion frameworks.
  • Understanding of data governance and data quality practices.
Preferred Candidate Profile
  • Strong analytical and problem‑solving skills.
  • Ability to diagnose and resolve complex Spark performance issues.
  • Excellent communication and stakeholder management skills.
  • Experience working in Agile/Scrum environments.
  • Self‑driven individual capable of working independently and within a team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
Spark Engineer
Spark Engineer

Impronics Technologies • Gurugram District

On-site
INR 1,200,000 - 1,800,000
PySpark Engineer
PySpark Engineer

Pagaar India • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Chennai District

On-site
INR 900,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Spark + Scala+ Python +Github Copilot
Spark + Scala+ Python +Github Copilot

SPG Consulting • Bengaluru Urban

On-site
INR 1,200,000 - 2,300,000