PySpark Developer - Data Engineering

Infosys

Bengaluru

On-site

INR 1,500,000 - 2,100,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Restaurant d'entreprise
Indemnités de stage/alternance

Job summary

Infosys is seeking a data engineer in Bengaluru to design, develop, and maintain scalable PySpark-based ETL/ELT pipelines for batch and incremental processing.

The role requires strong PySpark skills, experience with Spark concepts, and collaboration with cross-functional teams to ensure data quality and reliability.

Qualifications

  • Education: BTECH, MTECH, MCA, MSC.
  • 2–3 years of data engineering experience in large-scale processing.
  • Strong hands-on experience with PySpark for data pipelines.
  • Working knowledge of Spark concepts such as RDD/DataFrame, joins, shuffles, and performance.
  • Ability to debug Spark applications and resolve data/job issues.

Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines with PySpark.
  • Build and optimize Spark jobs focusing on partitioning, caching, and transformations.
  • Perform data cleansing, validation, and reconciliation for data quality.
  • Collaborate with cross-functional teams to translate requirements.
  • Troubleshoot pipeline failures and improve reliability and throughput.
  • Write clean, reusable code with documentation for pipelines.
  • Support deployment and production monitoring of Spark workloads.
  • Participate in code reviews and follow engineering best practices.

Skills

PySpark
SQL
Hadoop
Hive
Kafka
Airflow

Education

BTECH
MTECH
MCA
MSC

Job description

Pyspark Good to have skills: SQL, Hadoop, Hive, Kafka, Airflow

Preferred Qualifications
  • Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
  • Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
  • Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
  • Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.
Key Responsibilities
  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability.
Minimum Qualifications
  • Education: BTECH, MTECH, MCA, MSC.
  • 2–3 years of experience in data engineering or large-scale data processing roles.
  • Strong hands-on experience with PySpark for building data pipelines and transformations.
  • Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
  • Ability to debug Spark applications and resolve data/job issues effectively.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
Big Data Engineer - Hadoop & PySpark Lead
Big Data Engineer - Hadoop & PySpark Lead

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
PySpark Developer (2 To 3 Years)
PySpark Developer (2 To 3 Years)

Infosys • Dadri, Chennai District, Bengaluru

Hybrid
INR 600,000 - 900,000
PySpark, Spark Developer
PySpark, Spark Developer

Infosys • Bengaluru

On-site
INR 900,000 - 1,500,000
Hadoop / PySpark
Hadoop / PySpark

Infosys • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000
Python PySpark Developer
Python PySpark Developer

Hexaware Technologies • Hyderabad, Pune District, Bengaluru

On-site
INR 1,200,000 - 2,100,000