Senior PySpark ETL Lead Engineer

Relevantz Technology Services

Chennai District

On-site

INR 1,500,000 - 2,100,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Relevantz Technology Services is seeking a Senior PySpark ETL Engineer to design, build, and optimize scalable data pipelines using PySpark across enterprise data platforms.

You will own pipelines end-to-end in production, ensure data quality, performance, and cost efficiency, and collaborate with cross-functional teams on complex data challenges.

Qualifications

  • 10-12 years of IT experience with focus on data engineering and ETL
  • 3+ years hands-on PySpark/Apache Spark in production
  • Strong experience designing and implementing ETL/ELT pipelines at scale
  • Excellent knowledge of SQL and relational data concepts
  • Experience handling large datasets in distributed environments
  • Strong ownership mindset and ability to own production pipelines

Responsibilities

  • Design, implement, and maintain high-performance PySpark ETL pipelines
  • Own pipelines end-to-end, including development, deployment, monitoring, and production support
  • Ensure pipelines are scalable, fault-tolerant, and runnable
  • Implement incremental processing and efficient data movement strategies

Skills

PySpark
Spark SQL
ETL pipelines
SQL
Python
Data modeling
Batch processing
Airflow
Databricks
AWS EMR

Tools

Airflow
Databricks
AWS EMR
Kubernetes
S3

Job description

Position: Senior PySpark ETL Engineer

Role Summary

The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and operating scalable data pipelines using Apache Spark (PySpark). This role focuses on high volume batch (and optionally streaming) data processing, ensuring performance, reliability, data quality, and cost eciency across enterprise data platforms.

The position requires strong python, hands on Spark expertise, deep SQL and data modeling knowledge, and the ability to own pipelines end to end in production.

Mandatory Requirements
  • 10 - 12 years of overall IT experience, with strong focus on data engineering and ETL.
  • 3+ years of hands on experience with PySpark / Apache Spark in production environments.
  • Strong experience designing and implementing ETL / ELT pipelines at scale.
  • Excellent knowledge of SQL and relational data concepts.
  • Experience handling large datasets in distributed environments.
  • Strong ownership mindset, problem solving skills, and ability to independently handle production pipelines.
Core Technical Skills
PySpark & Spark Engineering
  • Deep expertise in PySpark:
    DataFrames, Spark SQL, window functions, joins, aggregations
  • Spark execution model (DAGs, stages, tasks)
  • Strong hands on experience with:
    Partitioning strategies
  • Shuffle optimization
  • Broadcast vs sort merge joins
  • Caching / persisting
  • Handling data skew and memory spills
  • Proven ability to debug and optimize slow Spark jobs.
ETL & Data Engineering:
  • Strong knowledge of ETL/ELT design patterns:
    Incremental loads
  • Watermarking
  • Idempotent pipeline design
  • Reprocessing and backfill strategies
  • Experience implementing:
    SCD Type 1 / Type 2
  • Deduplication and late arriving data handling
  • Ability to design reusable transformation frameworks and common utilities.
  • Experience building source to target reconciliation and data quality checks.
Data Storage & SQL
  • Excellent SQL skills including:
    Complex joins
  • Subqueries and CTEs
  • Window functions
  • Query optimization
  • Experience working with:
    RDBMS sources (Postgres, MySQL)
  • Data lake storage using Parquet / ORC
  • Experience with partitioned datasets and compaction strategies.
Cloud & Big Data Platforms
  • Hands on experience with at least one Spark platform:
    AWS EMR
  • Spark on Kubernetes
  • Experience working with cloud storage:
    S3
  • Familiarity with orchestration tools:
    Airflow, Databricks Workflows, ADF, or equivalent.
Responsibilities
Pipeline Development & Ownership
  • Design, implement, and maintain high performance PySpark ETL pipelines.
  • Own pipelines end to end, including development, deployment, monitoring, and production support.
  • Ensure pipelines are scalable, fault tolerant, and re runnable.
  • Implement incremental processing and ecient data movement strategies.
Performance & Reliability
  • Identify and fix Spark performance bottlenecks.
  • Optimize resource usage and reduce execution time and cost.
  • Handle production issues related to:
    Job failures
  • Data corruption
  • SLA breaches
  • Perform root cause analysis and implement permanent fixes.
Data Quality & Governance
  • Implement strong data quality validations, checks, and reconciliation mechanisms.
  • Ensure correctness, completeness, and freshness of datasets.
  • Follow enterprise standards for:
    Data retention
  • Auditability
  • Schema evolution
Engineering Excellence
  • Write clean, maintainable, and testable PySpark code.
  • Conduct code reviews and guide junior engineers.
  • Follow best practices for:
    Version control (Git)
  • CI/CD
  • Logging and monitoring
  • Maintain clear documentation and operational runbooks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Pyspark Developer
Pyspark Developer

Sightspectrum • Chennai District

On-site
INR 900,000 - 1,300,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
PySpark / Spark Developer
PySpark / Spark Developer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000
PySpark Engineer
PySpark Engineer

Pagaar India • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion Engineering • Bengaluru

On-site
INR 800,000 - 1,200,000
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000
Walk-in | Pyspark Data engineer
Walk-in | Pyspark Data engineer

Tata Consultancy Services • Pune District

On-site
INR 1,000,000 - 1,400,000