Position: Senior PySpark ETL Engineer
Role Summary
The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and operating scalable data pipelines using Apache Spark (PySpark). This role focuses on high volume batch (and optionally streaming) data processing, ensuring performance, reliability, data quality, and cost eciency across enterprise data platforms.
The position requires strong python, hands on Spark expertise, deep SQL and data modeling knowledge, and the ability to own pipelines end to end in production.
Mandatory Requirements
- 10 - 12 years of overall IT experience, with strong focus on data engineering and ETL.
- 3+ years of hands on experience with PySpark / Apache Spark in production environments.
- Strong experience designing and implementing ETL / ELT pipelines at scale.
- Excellent knowledge of SQL and relational data concepts.
- Experience handling large datasets in distributed environments.
- Strong ownership mindset, problem solving skills, and ability to independently handle production pipelines.
Core Technical Skills
PySpark & Spark Engineering
- Deep expertise in PySpark:
DataFrames, Spark SQL, window functions, joins, aggregations - Spark execution model (DAGs, stages, tasks)
- Strong hands on experience with:
Partitioning strategies - Shuffle optimization
- Broadcast vs sort merge joins
- Caching / persisting
- Handling data skew and memory spills
- Proven ability to debug and optimize slow Spark jobs.
ETL & Data Engineering:
- Strong knowledge of ETL/ELT design patterns:
Incremental loads - Watermarking
- Idempotent pipeline design
- Reprocessing and backfill strategies
- Experience implementing:
SCD Type 1 / Type 2 - Deduplication and late arriving data handling
- Ability to design reusable transformation frameworks and common utilities.
- Experience building source to target reconciliation and data quality checks.
Data Storage & SQL
- Excellent SQL skills including:
Complex joins - Subqueries and CTEs
- Window functions
- Query optimization
- Experience working with:
RDBMS sources (Postgres, MySQL) - Data lake storage using Parquet / ORC
- Experience with partitioned datasets and compaction strategies.
Cloud & Big Data Platforms
- Hands on experience with at least one Spark platform:
AWS EMR - Spark on Kubernetes
- Experience working with cloud storage:
S3 - Familiarity with orchestration tools:
Airflow, Databricks Workflows, ADF, or equivalent.
Responsibilities
Pipeline Development & Ownership
- Design, implement, and maintain high performance PySpark ETL pipelines.
- Own pipelines end to end, including development, deployment, monitoring, and production support.
- Ensure pipelines are scalable, fault tolerant, and re runnable.
- Implement incremental processing and ecient data movement strategies.
Performance & Reliability
- Identify and fix Spark performance bottlenecks.
- Optimize resource usage and reduce execution time and cost.
- Handle production issues related to:
Job failures - Data corruption
- SLA breaches
- Perform root cause analysis and implement permanent fixes.
Data Quality & Governance
- Implement strong data quality validations, checks, and reconciliation mechanisms.
- Ensure correctness, completeness, and freshness of datasets.
- Follow enterprise standards for:
Data retention - Auditability
- Schema evolution
Engineering Excellence
- Write clean, maintainable, and testable PySpark code.
- Conduct code reviews and guide junior engineers.
- Follow best practices for:
Version control (Git) - CI/CD
- Logging and monitoring
- Maintain clear documentation and operational runbooks.