Python-Pyspark Developer

Infosys

Hyderabad

On-site

INR 2,400,000 - 3,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infosys is seeking an experienced Python PySpark Developer to design, develop, and optimize large-scale data processing systems. You will work on big data platforms to build scalable ETL pipelines and process high-volume datasets using Spark and Python.

The role focuses on data engineering, distributed processing, and collaboration with data scientists and analysts to translate business requirements into technical solutions and ensure reliable data workflows.

Qualifications

  • Experience with Python and PySpark for large-scale data processing pipelines.
  • Proficient in designing and optimizing ETL/ELT workflows.
  • Strong understanding of big data processing on distributed systems.

Responsibilities

  • Develop data pipelines using Python and PySpark.
  • Process and transform large datasets in distributed environments.
  • Build scalable ETL/ELT workflows for batch and real-time processing.
  • Ingest data from databases, APIs, and files across platforms including Hadoop and cloud services.
  • Tune Spark jobs for performance and cost efficiency.
  • Collaborate with data engineers, data scientists, and analysts; participate in code reviews and agile practices.

Tools

Python
PySpark

Job description

  • Primary skills:Python, Pyspark

We are looking for an experienced Python PySpark Developer to design, develop, and optimize large-scale data processing systems. The ideal candidate will work on big data platforms, build scalable ETL pipelines, and process high-volume datasets using Spark and Python.

Key Responsibilities

Data Engineering & Development Develop and maintain data pipelines using Python and PySpark Process and transform large datasets in distributed environments Build scalable ETL/ELT workflows Big Data Processing Work with Apache Spark (PySpark) for batch and real-time processing Optimize Spark jobs for performance and efficiency Handle structured and unstructured data Data Integration Ingest data from multiple sources: Databases (SQL/NoSQL) APIs Files (CSV, JSON, Parquet) Integrate with data platforms like: Hadoop (HDFS) Cloud (AWS, Azure, GCP) Performance Optimization Tune Spark jobs (partitioning, caching, parallelism) Optimize SQL queries and transformations Improve data processing efficiency and cost Collaboration & Support Work with data engineers, data scientists, and analysts Translate business requirements into technical solutions Participate in code reviews and agile development practices Monitoring & Troubleshooting Debug and resolve issues in data pipelines Monitor job execution and data quality Ensure reliability and availability of data workflows

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Python, Pyspark Developer
Python, Pyspark Developer

Infosys • Hyderabad

On-site
INR 1,100,000 - 2,300,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer (Python & PySpark)
Data Engineer (Python & PySpark)

Techknomatic Services • Pune District

On-site
INR 700,000 - 1,200,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,200,000 - 1,500,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Python Data Engineer
Python Data Engineer

Lonvec Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Python, PySpark, ETL Developer
Python, PySpark, ETL Developer

Infosys • Hyderabad

On-site
INR 800,000 - 1,600,000
Python Pyspark Developer SIN
Python Pyspark Developer SIN

Tata Consultancy Services • Bengaluru

On-site
INR 1,200,000 - 1,600,000
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000