Pyspark

Infosys

Bengaluru

On-site

INR 1,500,000 - 2,500,000

Full time

9 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Infosys in Bengaluru seeks a data engineer to design and maintain scalable PySpark ETL/ELT pipelines for batch and incremental processing.

You will optimize Spark jobs, cleanse data, and collaborate with cross-functional teams to translate requirements into robust data solutions, while ensuring reliability and performance in production environments.

Qualifications

  • 2–3 years of experience in data engineering or large-scale data processing roles.
  • Strong hands-on experience with PySpark for building data pipelines and transformations.
  • Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
  • Ability to debug Spark applications and resolve data/job issues effectively.

Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability.

Skills

PySpark
Apache Spark
Hadoop
Hive
Kafka
Airflow

Education

BTECH/BE
MTECH/ME
MCA
MSC

Tools

Spark

Job description

Good to have skills: SQL, Hadoop, Hive, Kafka, Airflow

Key Responsibilities:
  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability.
Minimum Qualifications:
  • Education: BTECH, MTECH, MCA, MSC.
  • 2–3 years of experience in data engineering or large-scale data processing roles.
  • Strong hands‑on experience with PySpark for building data pipelines and transformations.
  • Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
  • Ability to debug Spark applications and resolve data/job issues effectively.
Preferred Qualifications:
  • Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
  • Exposure to building end‑to‑end data pipelines with strong data quality checks and automated validations.
  • Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
  • Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

PySpark Developer - Data Engineering
PySpark Developer - Data Engineering

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Restaurant d'entreprise
Indemnités de stage/alternance
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
Spark
Spark

Infosys • Bengaluru

On-site
INR 900,000 - 1,400,000
Hadoop / PySpark
Hadoop / PySpark

Infosys • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
Big Data Engineer - Hadoop & PySpark Lead
Big Data Engineer - Hadoop & PySpark Lead

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Databricks
Databricks

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,300,000
PySpark Developer (2 To 3 Years)
PySpark Developer (2 To 3 Years)

Infosys • Dadri, Chennai District, Bengaluru

Hybrid
INR 600,000 - 900,000