Developer - PySpark

Compunnel, Inc.

Pune District

On-site

INR 800,000 - 1,500,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading tech company in Pune is seeking skilled PySpark Developers to design and develop scalable data processing solutions. The role involves building ETL pipelines and collaborating with cross-functional teams to ensure data quality and performance. Candidates should have 2-10 years of experience in data engineering with a strong expertise in PySpark, Python, and big data platforms.

Qualifications

  • 2–10 years of experience in data engineering with strong hands-on expertise in PySpark.
  • Proficiency in Python programming for data manipulation and pipeline development.
  • Strong knowledge of Spark SQL, DataFrames, and RDDs.
  • Experience with big data ecosystems like Hadoop, Hive, or Databricks.
  • Good understanding of SQL and relational databases.
  • Familiarity with cloud platforms and services.
  • Experience with version control and CI/CD tools.

Responsibilities

  • Design, develop, and optimize ETL pipelines using PySpark.
  • Ingest, clean, transform, and process various data formats.
  • Work with large-scale datasets on distributed computing platforms.
  • Integrate data from multiple sources.
  • Ensure high-performance and reliability of data pipelines.
  • Collaborate with cross-functional teams.

Skills

PySpark
Python
SQL
Hadoop
Docker
Kubernetes
DataFrames
Git
Cloud platforms (AWS, Azure, GCP)

Education

Bachelor’s degree in computer science or computer engineering

Tools

Hadoop
Databricks
Docker
Kubernetes

Job description

We are looking for skilled PySparkDevelopers to design and develop scalable data processing solutions. The role involves working with big data platforms, building ETL pipelines, and collaborating with cross-functional teams to ensure data availability, performance, and quality.

Experience: 2-10 years (Mandatory)

Key Responsibilities
  • Design, develop, and optimize ETL/ELT pipelines using PySpark.
  • Ingest, clean, transform, and process structured/semi-structured/unstructured data.
  • Work with large-scale datasets on distributed computing platforms (Hadoop, Spark, Databricks).
  • Integrate data from multiple sources including Delta Lake, Data Lake, RDBMS, APIs.
  • Ensure high-performance, scalability, and reliability of data pipelines.
  • Collaborate with data engineers, analysts, and data scientists to deliver business-ready datasets.
  • Implement monitoring, logging, and error-handling for data pipelines.
  • Contribute to data quality frameworks and best practices.
Required Skills & Experience
  • Bachelor’s degree in computer science, computer engineering or similar.
  • 2–10 years of experience in data engineering with strong hands-on expertise in PySpark.
  • Proficiency in Python programming for data manipulation and pipeline development.
  • Strong knowledge of Spark SQL, DataFrames, and RDDs.
  • Experience with big data ecosystems – Hadoop, Hive, or Databricks.
  • Good understanding of SQL and relational databases (PostgreSQL, MySQL, Oracle).
  • Familiarity with cloud platforms (AWS, Azure, GCP) and services (S3, ADLS, BigQuery).
  • Experience with version control (Git) and CI/CD tools.
  • Strong debugging, performance optimization, and troubleshooting skills.
  • Experience with containerization (Docker, Kubernetes).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
PySpark Developer (2 To 3 Years)
PySpark Developer (2 To 3 Years)

Infosys • Dadri, Chennai District, Bengaluru

Hybrid
INR 600,000 - 900,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
PySpark Developer - Data Engineering
PySpark Developer - Data Engineering

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Restaurant d'entreprise
Indemnités de stage/alternance
Python PySpark Developer
Python PySpark Developer

Hexaware Technologies • Hyderabad, Pune District, Bengaluru

On-site
INR 1,200,000 - 2,100,000
PySpark Developer (2 To 8 Years)
PySpark Developer (2 To 8 Years)

Infosys • Dadri, Chennai District, Bengaluru

On-site
INR 900,000 - 1,800,000
Senior ETL / Data Engineer
Senior ETL / Data Engineer

Durapid Technologies Pvt Ltd • India

On-site
INR 2,500,000 - 4,500,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Contractor - PySpark Engineer
Contractor - PySpark Engineer

Vivantify • Hyderabad

On-site
INR 1,800,000 - 2,600,000