Python, PySpark, ETL Developer

Infosys

Hyderabad

On-site

INR 800,000 - 1,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infosys in Hyderabad seeks an experienced Data Pipeline Developer to build and maintain scalable ETL pipelines using Python and PySpark. You will ensure data quality, implement modular transformations, and optimize Spark jobs for performance and cost efficiency.

The role requires 2–5 years of hands-on experience with data pipelines, strong ETL knowledge, and a commitment to clean, maintainable code and best practices. Join a collaborative team delivering reliable data solutions.

Qualifications

  • Bachelors in CS/Engineering or related field, or equivalent practical experience.
  • 2–5 years of hands-on experience building data pipelines with Python and PySpark.
  • Strong understanding of ETL concepts, data transformations, and large datasets.
  • Proficient in writing clean, maintainable code and debugging production issues.
  • Knowledge of data structures, algorithms, and software development best practices.

Responsibilities

  • Develop and maintain scalable batch ETL pipelines using Python and PySpark.
  • Implement reusable transformation logic and modular, testable pipelines.
  • Optimize Spark jobs for performance and cost efficiency.
  • Apply data validation, handle schema evolution, ensure data accuracy.
  • Troubleshoot pipeline failures, analyze logs, and implement robust error handling.
  • Collaborate with cross-functional teams to define data mappings and deliver datasets.
  • Document pipeline logic and operational procedures for handovers.

Skills

Python
PySpark
ETL concepts
Data processing
Software development best practices

Education

Bachelor's degree in Computer Science, Engineering, Information Systems, or related field

Job description

Technology->Analytics - Packages->Python - Big Data,Technology->Big Data - Data Processing->PySpark, ETL

Data Pipeline Development
  • Develop and maintain scalable batch ETL pipelines using Python and PySpark for data ingestion, transformation, and loading.
  • Implement reusable transformation logic, ensuring pipelines are modular, testable, and easy to maintain.
  • Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and cost efficiency.
Data Quality & Reliability
  • Apply data validation checks, handle schema evolution, and ensure accuracy and completeness of processed datasets.
  • Troubleshoot pipeline failures, analyze logs, and implement robust error handling and retry mechanisms.
  • Monitor job runs and support operational stability through alerts, runbooks, and timely incident resolution.
Collaboration & Delivery
  • Work with cross-functional teams to gather requirements, define data mappings, and deliver datasets aligned to business needs.
  • Participate in code reviews, follow engineering best practices, and contribute to continuous improvement of standards and tooling.
  • Document pipeline logic, dependencies, and operational procedures for smooth handovers and long-term maintainability.
Qualifications
  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
  • 2–5 years of hands-on experience building data pipelines using Python and PySpark.
  • Strong understanding of ETL concepts, data transformations, and handling large-scale datasets.
  • Proficiency in writing clean, maintainable code and debugging production issues.
  • Working knowledge of data structures, algorithms, and software development best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Python, Pyspark Developer
Python, Pyspark Developer

Infosys • Hyderabad

On-site
INR 1,100,000 - 2,300,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion Engineering • Bengaluru

On-site
INR 800,000 - 1,200,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,200,000 - 1,500,000
Python with Data Engineer
Python with Data Engineer

Tata Consultancy Services • Chennai District, Bengaluru, Hyderabad

On-site
INR 1,800,000 - 3,000,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Python Data Engineer
Python Data Engineer

Lonvec Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Python_Data Engineering
Python_Data Engineering

Cognizant • Gurugram District

On-site
INR 600,000 - 1,200,000
Python-Pyspark Developer
Python-Pyspark Developer

Infosys • Hyderabad

On-site
INR 2,400,000 - 3,600,000
Senior PySpark ETL Lead Engineer
Senior PySpark ETL Lead Engineer

Relevantz Technology Services • Chennai District

On-site
INR 1,500,000 - 2,100,000