PySpark Engineer

Pagaar India

Hyderabad

On-site

INR 1,200,000 - 2,100,000

Full time

13 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pagaar India is seeking a PySpark Engineer to design, build, and optimise Spark applications and ETL data pipelines for a complex enterprise data lake. The role includes Delta Lake implementation, performance tuning, and CI/CD automation for PySpark, with emphasis on batch and streaming workloads.

You will work with Agile teams, size Spark clusters, and guide a small technical team to deliver high-availability Spark clusters and robust data strategies.

Qualifications

  • Proficient in designing and developing Spark applications using PySpark.
  • Hands-on experience building and maintaining ETL data pipelines.
  • Strong Python development skills.
  • Proficient in writing SQL scripts.
  • Experience with Spark Streaming and Structured Streaming.
  • Experience implementing Delta Lake in enterprise environments.
  • Knowledge of multiple cluster managers (Spark Standalone, YARN, Kubernetes).
  • Familiar with Big Data on Cloud, especially GCP Dataproc and GCS.

Responsibilities

  • Design and develop Spark applications for enterprise-scale data processing.
  • Build and maintain ETL pipelines feeding a complex data lake.
  • Implement Delta Lake for reliable data storage.
  • Tune and optimise Spark applications for performance.
  • Develop Spark Streaming workloads.
  • Set up CI/CD pipelines for PySpark deployments.
  • Write and run test cases, including performance tests.
  • Size clusters and manage resources across Spark Standalone, YARN, and Kubernetes.
  • Provide technical guidance and lead a small team of specialists.

Skills

PySpark
ETL pipelines
Python
SQL
Spark Streaming
Delta Lake
Cluster management
GCP Dataproc
CI/CD
Testing Spark apps

Tools

Airflow
Cloud Composer

Job description

Role Summary

We are seeking a PySpark Engineer to design, build, and optimise Spark applications and ETL data pipelines for a complex enterprise data lake. The role covers batch and streaming workloads, Delta Lake implementation, Spark performance tuning and cluster sizing, and CI/CD automation for PySpark. The engineer will operate highly available Spark clusters with monitoring, work with Agile application development teams on data strategy and dataflows, and act as a technical specialist with responsibility for guiding a small team.

Key Responsibilities
  • Design and develop Spark applications using PySpark for enterprise-scale data processing.
  • Build and maintain ETL data pipelines feeding a complex data lake implementation.
  • Implement Delta Lake (delta.io) for enterprise-grade data storage and reliability.
  • Perform optimisation and performance tuning of Spark applications.
  • Develop Spark Streaming and Structured Streaming workloads.
  • Build and set up CI/CD pipelines for PySpark deployments.
  • Write and execute test cases for Spark applications, including performance tests.
  • Size Spark clusters and manage resources across Spark Standalone, YARN, and Kubernetes cluster managers.
  • Set up and operate highly available Spark clusters with operational monitoring.
  • Work with Agile application development teams to implement data strategies, build dataflows, and define conceptual data models.
  • Forecast environment requirements based on anticipated demand from multiple application development teams.
  • Create short-term plans to deliver environments supporting sprint-based development.
  • Provide technical guidance and manage a small team of technical specialists.
Primary Skills (Must Have)
  • Experience designing and developing Spark applications using PySpark.
  • Hands-on experience building and maintaining ETL data pipelines.
  • Expertise in Python development.
  • Proficiency in writing SQL scripts.
  • Spark Streaming and Structured Streaming knowledge is mandatory.
  • Experience with Delta Lake (delta.io) for enterprise-grade implementation.
  • Optimisation and performance tuning of Spark applications.
  • Experience with different cluster managers:
    • Spark Standalone
    • YARN
    • Kubernetes
  • Spark cluster sizing and resource management for a complex data lake implementation.
  • Experience setting up and operating highly available Spark clusters with operational monitoring.
  • Experience building and setting up CI/CD pipelines for PySpark.
  • Experience writing and executing test cases for Spark applications, including performance tests.
  • Knowledge of Big Data on Cloud, preferably GCP services such as Dataproc and GCS.
  • Strong communication skills and the ability to plan and prioritise own time effectively.
  • Ability to manage a small team as technical specialists.
Secondary Skills (Nice to Have)
  • Java or Scala development experience.
  • AWS or Azure cloud platform experience.
  • Exposure to workflow orchestration tools such as Airflow or Cloud Composer.
  • Familiarity with data governance, lineage, and cataloguing practices.
  • Observability tooling for Spark workloads metrics, logging, and alerting.

GCP Data Engineer certification is an advantage

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Contractor - PySpark Engineer
Contractor - PySpark Engineer

Vivantify • Hyderabad

On-site
INR 1,800,000 - 2,600,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Analytics Engineer
Data Analytics Engineer

Monocept • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior PySpark ETL Lead Engineer
Senior PySpark ETL Lead Engineer

Relevantz Technology Services • Chennai District

On-site
INR 1,500,000 - 2,100,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000
Lead Data Engineer
Lead Data Engineer

ORMAE • Pune District

On-site
INR 1,500,000 - 2,100,000