Hadoop / PySpark

Infosys

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Infosys in Bengaluru seeks a Senior Lead Data Engineer to drive end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.

You will design scalable ETL pipelines, optimize Spark jobs, mentor engineers, and establish data quality checks and runbooks to support production readiness.

Qualifications

  • Bachelor's or Master's degree in Engineering/Technology/Science or equivalent.
  • 9–11 years of experience in data engineering and large-scale processing.
  • Strong hands-on experience with Hadoop ecosystem and PySpark.

Responsibilities

  • Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.
  • Design and implement scalable batch/ETL pipelines for large datasets with reliability and performance focus.
  • Drive architecture discussions, define standards, and ensure best practices for distributed data processing.
  • Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.
  • Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.
  • Perform root-cause analysis for pipeline failures and implement preventive fixes.

Skills

Leadership
Stakeholder management
PySpark
Hadoop
Data modeling
Distributed processing
Code reviews
Mentoring
Consulting mindset

Education

BE/BTech/MTech/MCA/MSc or equivalent

Tools

Hadoop
PySpark
Spark

Job description

  • Primary skills:Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop
  • Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.
  • Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance.
  • Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing.
  • Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable.
  • Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones.
  • Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.
  • Perform root-cause analysis for pipeline failures and performance bottlenecks; implement preventive fixes.
  • Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.
  • Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness.
  • Bachelor’s or Master’s degree (or equivalent) in Engineering/Technology/Computer Applications/Science (BE/BTech/MTech/MCA/MSc or equivalent).
  • 9–11 years of overall experience in data engineering and large-scale data processing environments.
  • Strong hands-on experience with Hadoop ecosystem components and distributed data processing concepts.
  • Strong hands-on experience building data pipelines using PySpark.
  • Proven ability to lead technical delivery, guide teams, and manage stakeholder expectations in consulting engagements. Preferred Qualifications:
  • Experience designing and implementing Spark-based data processing patterns (batch and incremental loads).
  • Strong understanding of data modeling and storage patterns for big data platforms (partitioning, compaction, schema evolution).
  • Experience with workflow orchestration and scheduling for data pipelines and dependency management.
  • Demonstrated expertise in production hardening: monitoring, alerting, SLAs, and incident management for data jobs.
  • Strong consulting mindset with ability to present solutions, document designs clearly, and influence technical decisions across teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Big Data Engineer - Hadoop & PySpark Lead
Big Data Engineer - Hadoop & PySpark Lead

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,600,000
PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
PySpark Developer - Data Engineering
PySpark Developer - Data Engineering

Infosys • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Restaurant d'entreprise
Indemnités de stage/alternance
Big Data Engineer
Big Data Engineer

Nice Software Solutions Pvt. Ltd. • Pune District

On-site
INR 1,800,000 - 3,000,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Senior ETL / Data Engineer
Senior ETL / Data Engineer

Durapid Technologies Pvt Ltd • India

On-site
INR 2,500,000 - 4,500,000
Big Data Engineer Lead
Big Data Engineer Lead

Evoke HR • Chennai District, Bengaluru

On-site
INR 2,500,000 - 4,000,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000
Data Engineer-Pyspark
Data Engineer-Pyspark

Deloitte US-India Offices • Chennai District

On-site
INR 900,000 - 1,300,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion • Hyderabad

On-site
INR 2,400,000 - 4,200,000