PySpark Developer (2 To 3 Years)

Infosys

Dadri, Chennai District, Bengaluru

Hybrid

INR 600,000 - 900,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Infosys is seeking a PySpark Developer to design, implement, and optimize large-scale data pipelines using PySpark and Spark SQL. You will work on ETL/ELT processes, transform data with DataFrames, and ensure data quality across distributed systems.

The role requires strong Python programming, SQL proficiency, and hands-on Spark experience in a PAN India context. The ideal candidate will collaborate with Data Engineers and Analysts to meet analytics requirements, monitor Spark jobs, and

Qualifications

  • 23 years of experience in PySpark development.
  • Strong programming skills in Python.
  • Hands-on experience with Apache Spark, Spark SQL, DataFrames, and RDDs.
  • Experience in ETL development and data integration projects.
  • Good understanding of data warehousing concepts.
  • Experience with relational databases such as SQL Server, Oracle, MySQL, or PostgreSQL.
  • Strong SQL querying and performance tuning skills.
  • Knowledge of Linux/Unix environments and shell scripting.
  • Understanding of version control systems such as Git.
  • Strong analytical and problem-solving skills.

Responsibilities

  • Design, develop, and maintain scalable data processing applications using PySpark.
  • Build and optimize ETL/ELT pipelines for structured and unstructured data.
  • Develop data transformation logic using Spark SQL and DataFrames.
  • Work with large-scale datasets in distributed computing environments.
  • Collaborate with Data Engineers, Data Analysts, and Business stakeholders to understand data requirements.
  • Monitor, troubleshoot, and optimize Spark jobs for performance and reliability.
  • Ensure data quality, integrity, and consistency across data pipelines.
  • Participate in code reviews and follow coding best practices.
  • Support deployments, enhancements, and production issue resolution.
  • Create technical documentation and maintain operational procedures.

Skills

PySpark development
Python
Spark SQL
DataFrames
RDDs
ETL development
Data warehousing
SQL databases
Git
Linux/Unix shell
Analytical thinking
Problem solving

Tools

Apache Spark
SQL Server
Oracle
MySQL
PostgreSQL
Git

Job description

Job Title: PySpark Developer
Experience: 2-3 Years
Location: PAN India
Employment Type: Full-Time

Job Summary

We are seeking a skilled PySpark Developer with 23 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have hands-on experience with PySpark, Python, Spark SQL, ETL development, and Big Data technologies. The role involves building scalable data pipelines, transforming large datasets, and supporting analytics and reporting requirements.

Key Responsibilities
  • Design, develop, and maintain scalable data processing applications using PySpark.
  • Build and optimize ETL/ELT pipelines for structured and unstructured data.
  • Develop data transformation logic using Spark SQL and DataFrames.
  • Work with large-scale datasets in distributed computing environments.
  • Collaborate with Data Engineers, Data Analysts, and Business stakeholders to understand data requirements.
  • Monitor, troubleshoot, and optimize Spark jobs for performance and reliability.
  • Ensure data quality, integrity, and consistency across data pipelines.
  • Participate in code reviews and follow coding best practices.
  • Support deployments, enhancements, and production issue resolution.
  • Create technical documentation and maintain operational procedures.

Required Skillsa
  • 23 years of experience in PySpark development.
  • Strong programming skills in Python.
  • Hands-on experience with Apache Spark, Spark SQL, DataFrames, and RDDs.
  • Experience in ETL development and data integration projects.
  • Good understanding of data warehousing concepts.
  • Experience working with relational databases such as SQL Server, Oracle, MySQL, or PostgreSQL.
  • Strong SQL querying and performance tuning skills.
  • Knowledge of Linux/Unix environments and shell scripting.
  • Understanding of version control systems such as Git.
  • Strong analytical and problem-solving skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Pyspark Developer
Pyspark Developer

Leading Global Technology Services Company • Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Contractor - PySpark Engineer
Contractor - PySpark Engineer

Vivantify • Hyderabad

On-site
INR 1,800,000 - 2,600,000
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
Walk-in | Pyspark Developer
Walk-in | Pyspark Developer

Tata Consultancy Services • Chennai District

On-site
INR 1,400,000 - 2,000,000
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000
Python, Pyspark Developer
Python, Pyspark Developer

Infosys • Hyderabad

On-site
INR 1,100,000 - 2,300,000