Python, PySpark, ETL Developer

Infosys

Hyderabad

On-site

INR 2,000,000 - 4,200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Infosys, a leader in digital services, is seeking a Data Analytics professional for its Data Analytics Unit in Hyderabad to design and build batch ETL pipelines using Python and PySpark.

You will implement modular transformations, optimize Spark jobs, and ensure data quality while collaborating with cross-functional teams and documenting procedures.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
  • 25 years of hands-on experience building data pipelines using Python and PySpark.
  • Strong understanding of ETL concepts, data transformations, and handling large-scale datasets.
  • Proficiency in writing clean, maintainable code and debugging production issues.
  • Working knowledge of data structures, algorithms, and software development best practices.

Responsibilities

  • Develop and maintain scalable batch ETL pipelines using Python and PySpark.
  • Implement reusable transformation logic for modular, testable pipelines.
  • Optimize Spark jobs for performance and cost efficiency.
  • Apply data validation, handle schema evolution, ensure data accuracy.
  • Troubleshoot pipeline failures, analyze logs, and implement robust error handling.
  • Monitor job runs and support operational stability with alerts and runbooks.
  • Collaborate with cross-functional teams to deliver datasets aligned to business needs.
  • Participate in code reviews and contribute to tooling standards.
  • Document pipeline logic, dependencies, and handover procedures.

Skills

Python
PySpark
Data Pipelines
ETL
Big Data
Data Quality

Education

Bachelor of Engineering

Tools

PySpark
Spark

Job description

Educational Requirements

Bachelor of Engineering

Service Line

Data Analytics Unit

Responsibilities
  • Data Pipeline DevelopmentDevelop and maintain scalable batch ETL pipelines using Python and PySpark for data ingestion, transformation, and loading.
  • Implement reusable transformation logic, ensuring pipelines are modular, testable, and easy to maintain.
  • Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and cost efficiency.
  • Data Quality Reliability: Apply data validation checks, handle schema evolution, and ensure accuracy and completeness of processed datasets.
  • Troubleshoot pipeline failures, analyze logs, and implement robust error handling and retry mechanisms.
  • Monitor job runs and support operational stability through alerts, runbooks, and timely incident resolution.
  • Collaboration Delivery: Work with cross-functional teams to gather requirements, define data mappings, and deliver datasets aligned to business needs.
  • Participate in code reviews, follow engineering best practices, and contribute to continuous improvement of standards and tooling.
  • Document pipeline logic, dependencies, and operational procedures for smooth handovers and long-term maintainability.
Additional Responsibilities
  • Bachelors degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
  • 25 years of hands-on experience building data pipelines using Python and PySpark.
  • Strong understanding of ETL concepts, data transformations, and handling large-scale datasets.
  • Proficiency in writing clean, maintainable code and debugging production issues.
  • Working knowledge of data structures, algorithms, and software development best practices.
Technical and Professional Requirements

Technology - Analytics - Packages - Python - Big Data, Technology - Big Data - Data Processing - PySpark, ETL

Preferred Skills
  • Technology - Analytics - Packages - Python - Big Data
  • Technology - Big Data - Data Processing - PySpark
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Python-Pyspark Developer
Python-Pyspark Developer

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Python, Spark Scala Developer
Python, Spark Scala Developer

Infosys • Hyderabad

On-site
INR 1,000,000 - 1,500,000
PySpark Databricks Engineer
PySpark Databricks Engineer

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
PySpark Developer
PySpark Developer

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
Python with Data Engineer
Python with Data Engineer

Tata Consultancy Services • Chennai District, Bengaluru, Hyderabad

On-site
INR 1,800,000 - 3,000,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion Engineering • Bengaluru

On-site
INR 800,000 - 1,200,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000