Senior Data Platform Engineer (Python/Spark)

Genpact

Bengaluru

On-site

INR 4,000,000 - 7,500,000

Full time

24 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Genpact in Bengaluru is seeking a Senior Data Platform Engineer with 10+ years of hands-on data engineering experience. You will design and evolve data ingestion pipelines using Python and Spark, and lead modernization initiatives leveraging cloud-native technologies.

The role requires deep knowledge of Hadoop ecosystem components like HDFS, Hive, and Cloudera-based deployments, plus experience with Kubernetes-based workloads and CI/CD pipelines.

Qualifications

  • 10+ years of hands-on data engineering and platform development experience.
  • Strong Python and Apache Spark development and framework design.
  • Experience building reusable and scalable data ingestion and transformation frameworks.
  • Deep understanding of Hive, Impala, HDFS, and the Hadoop ecosystem (preferably Cloudera-based).
  • Strong knowledge of HDFS internals, data storage architecture, partitioning, and performance optimization.
  • Ability to troubleshoot complex platform, storage, compute, and cluster issues and drive improvements.

Responsibilities

  • Design, develop, and maintain scalable data ingestion and transformation pipelines.
  • Analyze existing frameworks and implement new features and enhancements.
  • Drive platform modernization through cloud-native technologies, automation, and best practices.
  • Develop optimized integrations with object storage, databases, and enterprise data systems.
  • Build and support high-performance data pipelines on large-scale big data platforms.

Skills

Python
Apache Spark
Shell Scripting
Hadoop Ecosystem
HDFS Internals
Hive
Impala
Cloudera
Kubernetes
Databricks
JFrog
CI/CD
Cloud Data Platforms
AI Agents/LLM tools

Tools

JFrog Artifactory
Databricks
Cloudera Manager

Job description

Immediate joiners or candidates with 30-day notice will be preferred

We are looking for a highly skilled and self-driven Senior Data Platform Engineer with strong expertise in Python, Spark, Big Data Platforms, and Cloud Data Engineering. The ideal candidate should be capable of independently understanding existing data frameworks, enhancing them with new capabilities, and leading modernization initiatives involving cloud-native technologies.

Key Responsibilities :
  • Design, develop, and maintain scalable data ingestion and transformation frameworks using Python, Spark, and Shell Scripting.
  • Independently analyze existing frameworks and implement new features and enhancements.
  • Drive platform modernization through cloud-native technologies, automation, and engineering best practices.
  • Develop optimized integrations with object storage platforms, databases, and enterprise data systems.
  • Build and support high-performance data pipelines on large-scale big data platforms.
Required Skills :
  • 10 + Years of experience of hands on experience
  • Strong hands-on expertise in Python and Apache Spark development and framework design.
  • Experience building reusable and scalable data ingestion and transformation frameworks.
  • Deep understanding of Hive, Impala, HDFS, and the Hadoop ecosystem, preferably on Cloudera-based platforms.
  • Strong knowledge of HDFS internals, data storage architecture, compaction strategies, partition management, resource utilization, and cluster performance optimization.
  • Proven ability to troubleshoot complex platform, storage, compute, and cluster-level issues, perform root cause analysis, and drive performance improvements.
  • Experience working with diverse database technologies and object storage platforms, with a focus on optimized connectivity and data access patterns.
  • Strong understanding of distributed data processing, Spark optimization, and platform engineering best practices.
Preferred Skills :
  • Strong hands-on experience with Kubernetes-based application development, containerized Spark workloads, pod management, scaling, troubleshooting, and operational support.
  • Experience building and deploying Spark workloads using JFrog-managed container images and CI/CD pipelines.
  • Hands-on experience with Databricks and modern cloud-native data platforms.
  • Exposure to cloud migration and enterprise data platform modernization initiatives.
  • Experience leveraging AI Agents, Generative AI, and LLM-powered development tools to accelerate software delivery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Walk-in | Pyspark Developer
Walk-in | Pyspark Developer

Tata Consultancy Services • Chennai District

On-site
INR 1,400,000 - 2,000,000
Senior Data Engineer – PySpark, Cloud & Kafka - 5+ YoE - Immediate Joiner - Any UST Location
Senior Data Engineer – PySpark, Cloud & Kafka - 5+ YoE - Immediate Joiner - Any UST Location

UST • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,500,000 - 2,100,000
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)

Tata Consultancy Services • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,600,000
Lead Data Engineer (Databricks, PySpark & GCP)
Lead Data Engineer (Databricks, PySpark & GCP)

Egen • Hyderabad

On-site
INR 5,500,000 - 7,500,000
Healthcare benefits
Performance bonus
Data Engineer (AWS, Databricks, PySpark)
Data Engineer (AWS, Databricks, PySpark)

Tata Consultancy Services • Hyderabad, Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior Data Engineer
Senior Data Engineer

AagatiServe Pvt Ltd • Delhi

On-site
INR 1,800,000 - 2,400,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Mumbai, Bengaluru, New Delhi

On-site
INR 1,800,000 - 3,200,000
Senior Data Engineer
Senior Data Engineer

People, Jobs, and News • Dadri

On-site
INR 1,500,000 - 2,500,000
Data Engineer (PySpark + Cloudera)
Data Engineer (PySpark + Cloudera)

Zorba AI • Maharashtra

On-site
INR 1,000,000 - 1,500,000