Pyspark

Infosys

Bengaluru

On-site

INR 1,200,000 - 1,800,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Infosys is seeking a Data Engineer to design and optimize large-scale PySpark data pipelines in Bengaluru. You will build scalable ETL/ELT processes, ensure data quality, and collaborate with cross-functional teams to deliver reliable analytics-ready datasets.

The role focuses on hands-on development, performance tuning of Spark jobs, and contributing to robust data solutions in a fast-paced environment. You will work with modern big data technologies and grow your impact within the Data

Qualifications

  • Experience with PySpark data pipelines and transformations.
  • Experience with large-scale data processing.
  • Degree in engineering or computer science as listed (B.Tech/BE, M.Tech, MCA, MSc).
  • Ability to debug Spark applications and resolve data/job issues.

Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability.

Education

B.Tech/BE
M.Tech
MCA
MSc

Tools

PySpark
SQL
Hadoop
Hive
Kafka
Airflow

Job description

Job DescriptionAbout the job:

Join a collaborative data engineering team where your work directly powers reliable analytics and smarter business decisions. In this role, youll build and optimize scalable data processing pipelines using PySpark and Spark, working closely with engineers, analysts, and stakeholders to turn raw data into trusted, high-quality datasets. Youll be encouraged to take ownership, suggest improvements, and contribute to a culture that values clean engineering, performance, and continuous learning. If you enjoy solving data challenges, tuning distributed jobs, and delivering dependable solutions in a fast-moving environment, this is a great opportunity to grow your impact while working with modern big data technologies.

Roles ResponsibilitiesKey Responsibilities:
  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability.
Minimum Qualifications:
  • Education: BTECH, MTECH, MCA, MSC.
  • 23 years of experience in data engineering or large-scale data processing roles.
  • Strong hands-on experience with PySpark for building data pipelines and transformations.
  • Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
  • Ability to debug Spark applications and resolve data/job issues effectively.
Technical RequirementGood to have skills:

SQL, Hadoop, Hive, Kafka, Airflow

Additional ResponsibilityPreferred Qualifications:
  • Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
  • Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
  • Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
  • Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.
Educational RequirementMCA,MSc,MTech,Bachelor of Engineering,BTech
Preferred SkillsTechnology->Big Data - Data Processing->PySpark
Service LineData Analytics Unit
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

PySpark Developer - Data Engineering
PySpark Developer - Data Engineering

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Restaurant d'entreprise
Indemnités de stage/alternance
Spark-Scala, Databricks
Spark-Scala, Databricks

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Databricks
Databricks

Infosys • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Hadoop / PySpark
Hadoop / PySpark

Infosys • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Pyspark
Pyspark

Tata Consultancy Services • Chennai District

On-site
INR 1,800,000 - 2,600,000
Spark
Spark

Infosys • Bengaluru

On-site
INR 900,000 - 1,400,000
Python, PySpark, ETL Developer
Python, PySpark, ETL Developer

Infosys • Hyderabad

On-site
INR 2,000,000 - 4,200,000
Pyspark Developer
Pyspark Developer

Infosys • Hyderabad

On-site
INR 1,500,000 - 2,400,000
Pyspark Developer
Pyspark Developer

Infosys • Pune District

On-site
INR 1,200,000 - 2,100,000
Continuous learning programs
Certifications and career growth
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000