Role Summary
We are seeking a PySpark Engineer to design, build, and optimise Spark applications and ETL data pipelines for a complex enterprise data lake. The role covers batch and streaming workloads, Delta Lake implementation, Spark performance tuning and cluster sizing, and CI/CD automation for PySpark. The engineer will operate highly available Spark clusters with monitoring, work with Agile application development teams on data strategy and dataflows, and act as a technical specialist with responsibility for guiding a small team.
Key Responsibilities
- Design and develop Spark applications using PySpark for enterprise-scale data processing.
- Build and maintain ETL data pipelines feeding a complex data lake implementation.
- Implement Delta Lake (delta.io) for enterprise-grade data storage and reliability.
- Perform optimisation and performance tuning of Spark applications.
- Develop Spark Streaming and Structured Streaming workloads.
- Build and set up CI/CD pipelines for PySpark deployments.
- Write and execute test cases for Spark applications, including performance tests.
- Size Spark clusters and manage resources across Spark Standalone, YARN, and Kubernetes cluster managers.
- Set up and operate highly available Spark clusters with operational monitoring.
- Work with Agile application development teams to implement data strategies, build dataflows, and define conceptual data models.
- Forecast environment requirements based on anticipated demand from multiple application development teams.
- Create short-term plans to deliver environments supporting sprint-based development.
- Provide technical guidance and manage a small team of technical specialists.
Primary Skills (Must Have)
- Experience designing and developing Spark applications using PySpark.
- Hands-on experience building and maintaining ETL data pipelines.
- Expertise in Python development.
- Proficiency in writing SQL scripts.
- Spark Streaming and Structured Streaming knowledge is mandatory.
- Experience with Delta Lake (delta.io) for enterprise-grade implementation.
- Optimisation and performance tuning of Spark applications.
- Experience with different cluster managers:
- Spark Standalone
- YARN
- Kubernetes
- Spark cluster sizing and resource management for a complex data lake implementation.
- Experience setting up and operating highly available Spark clusters with operational monitoring.
- Experience building and setting up CI/CD pipelines for PySpark.
- Experience writing and executing test cases for Spark applications, including performance tests.
- Knowledge of Big Data on Cloud, preferably GCP services such as Dataproc and GCS.
- Strong communication skills and the ability to plan and prioritise own time effectively.
- Ability to manage a small team as technical specialists.
Secondary Skills (Nice to Have)
- Java or Scala development experience.
- AWS or Azure cloud platform experience.
- Exposure to workflow orchestration tools such as Airflow or Cloud Composer.
- Familiarity with data governance, lineage, and cataloguing practices.
- Observability tooling for Spark workloads metrics, logging, and alerting.
GCP Data Engineer certification is an advantage