Senior Spark / PySpark Data Engineer

Web Spiders

Kolkata Metropolitan Area

On-site

INR 1,200,000 - 1,800,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering.

The ideal candidate will design and develop high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, with cloud AWS services. Working hours align with US EST, office in Kolkata, immediate joiners preferred.

Qualifications

  • 5+ years of hands-on data engineering experience.
  • Strong hands-on Apache Spark experience.
  • Strong PySpark and Python programming skills.
  • Proven experience developing and optimizing large-scale ETL/ELT pipelines.
  • Experience with cloud-based data platforms, AWS services, and data lakes.

Responsibilities

  • Design, develop, and optimize large-scale ETL/ELT pipelines using Spark and PySpark.
  • Develop scalable data transformation and processing solutions with PySpark and Python.
  • Build distributed data-processing applications capable of handling large data volumes.
  • Develop reusable Spark/PySpark data-processing components.
  • Optimize Spark jobs for performance, scalability, and memory usage.
  • Tackle complex transformations, joins, aggregations, and partitioning.
  • Implement data validation, quality checks, error handling, and monitoring within pipelines.
  • Work with AWS data services such as EMR, Glue, S3, and Redshift.
  • Develop data pipelines supporting data lakes, warehouses, analytics, and downstream apps.
  • Troubleshoot production data pipelines and Spark processing issues.
  • Identify and resolve Spark/PySpark performance bottlenecks.
  • Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.

Skills

Spark
PySpark
Python
ETL/ELT
Distributed computing
Large datasets
Spark performance tuning
AWS
S3
Redshift
Glue
EMR
Airflow
MWAA
Data lakes

Tools

AWS EMR
AWS Glue
Amazon S3
Amazon Redshift
Airflow / MWAA
Step Functions
Hadoop ecosystem

Job description

Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering.

The ideal candidate will have strong experience designing and developing high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, along with experience working with cloud-based data platforms and AWS services.

If Spark + PySpark + Python is your core expertise and you enjoy solving complex data-processing and scalability challenges, we'd love to hear from you.

5+ Years Experience | Kolkata – Work from Office

Core Stack: Apache Spark

  • PySpark
  • Python
  • ETL/ELT
  • AWS
  • S3
  • Glue
  • Redshift
  • Immediate joiners preferred.*
Working Hours

Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.

What You'll Do
  • Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.
  • Develop scalable data transformation and processing solutions using PySpark and Python.
  • Build distributed data-processing applications capable of handling large volumes of data.
  • Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.
  • Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.
  • Work with complex transformations, joins, aggregations, partitioning, and large datasets.
  • Implement data validation, quality checks, error handling, and monitoring within data pipelines.
  • Work with AWS data services including EMR, Glue, S3, and Redshift.
  • Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
  • Troubleshoot production data pipeline and Spark processing issues.
  • Identify and resolve performance bottlenecks in Spark/PySpark workloads.
  • Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.
Must-Have Skills
  • 5+ years of hands-on experience in Data Engineering.
  • Strong hands-on experience with Apache Spark.
  • Strong hands-on experience with PySpark.
  • Strong programming experience in Python.
  • Proven experience developing and optimizing large-scale ETL/ELT pipelines.
  • Strong understanding of distributed computing and data-processing concepts.
  • Experience working with large datasets and complex data transformations.
  • Strong understanding of Spark performance optimization and tuning.
  • Experience with cloud-based data engineering, preferably AWS.
  • Experience with Amazon S3 and at least one AWS data-processing service such as EMR or Glue.
Good To Have
  • AWS EMR
  • AWS Glue
  • Apache Airflow / MWAA
  • AWS Step Functions
  • Amazon Redshift
  • Hadoop ecosystem
  • Experience with data lake and data warehouse architectures.
  • Experience with CI/CD and production deployment of data pipelines.
  • AWS Certified Data Engineer or another relevant AWS certification.
Interview Process
  • Application review
  • 5–10 minute initial screening call with the TA team
  • Technical interviews {Domain specific}
  • Practical test conducted in the presence of a panel member
  • Role match & offer
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AWS EMR Engineer
Senior AWS EMR Engineer

Web Spiders • Kolkata Metropolitan Area

On-site
INR 2,500,000 - 4,200,000
AWS + Pyspark Data Engineer ( Gurugram)
AWS + Pyspark Data Engineer ( Gurugram)

PwC India • Gurugram District

Hybrid
INR 1,400,000 - 2,100,000
Data Engineer — PySpark + AWS Glue
Data Engineer — PySpark + AWS Glue

Yadimen Consulting Limited • Chennai District, Pune District, Bengaluru

On-site
INR 1,000,000 - 1,800,000
Data Engineer - ETL
Data Engineer - ETL

Forward Eye Technologies • Pune District

On-site
INR 2,000,000 - 3,200,000
Senior Data Engineer - Pyspark & AWS
Senior Data Engineer - Pyspark & AWS

RBM Software • Pune District

On-site
INR 4,000,000 - 7,000,000
Pyspark Data Engineer
Pyspark Data Engineer

Tata Consultancy Services • Hyderabad, Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Dadri

Hybrid
INR 1,400,000 - 2,000,000
AWS & Pyspark- Data Engineer-Immediate joiners only
AWS & Pyspark- Data Engineer-Immediate joiners only

EY • Pune District, Bengaluru, Delhi

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer
Senior Data Engineer

AagatiServe Pvt Ltd • Delhi

On-site
INR 1,800,000 - 2,400,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,200,000 - 2,400,000