Data Architect - Pyspark

iSystems Ltd

Punjab

On-site

PKR 3,500,000 - 5,500,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

iSystems Ltd is seeking a Data Architect with strong hands-on PySpark experience to design, develop, optimize, and maintain scalable data processing solutions. You will work with large-scale data processing and distributed computing environments to deliver production-grade PySpark code.

The role requires 8+ years in Data Engineering, expertise in PySpark, Spark architecture, DataFrames, and cloud/on-premise data pipelines.

Qualifications

  • 8+ years of experience in Data Engineering or related field.
  • Strong and mandatory hands-on experience with PySpark.
  • Proven experience writing production-grade PySpark code.
  • Strong understanding of Spark architecture, DataFrames, Spark SQL, transformations, actions, partitioning, caching, and optimization.
  • Strong proficiency in Python with practical data engineering application development.
  • Hands-on experience building and managing large-scale ETL/ELT data pipelines.
  • Strong SQL skills with complex queries, joins, aggregations, and performance optimization.
  • Experience with cloud data platforms such as Azure, AWS, or GCP is preferred.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using PySpark.
  • Write clean, efficient, and production-ready PySpark code for large-scale data processing.
  • Develop ETL/ELT pipelines to ingest, transform, validate, and load data from multiple sources.
  • Perform complex data transformations, aggregations, joins, filtering, and data cleansing using PySpark.
  • Optimize PySpark jobs for performance, scalability, memory utilization, and execution time.
  • Work with distributed data processing frameworks and large datasets in cloud or on-premise environments.
  • Develop reusable PySpark frameworks, libraries, and data processing components.
  • Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues.
  • Implement data validation, error handling, logging, and monitoring within data pipelines.
  • Work closely with Data Architects, Data Engineers, Data Scientists, BI teams, and business stakeholders to understand data requirements.
  • Participate in data modeling, pipeline architecture, and technical design discussions.
  • Review code and provide technical guidance and mentorship to junior and mid-level Data Engineers.
  • Ensure adherence to coding standards, data engineering best practices, security, and governance requirements.
  • Work with CI/CD processes and version control systems for deploying and managing data pipelines.
  • Contribute to technical documentation and maintain clear documentation of data pipelines and processes.

Skills

PySpark
Spark Architecture
Python
SQL
ETL/ELT
CI/CD
Distributed Computing
Databricks
Delta Lake
Kafka
Hive
Git

Tools

Databricks
Delta Lake
Kafka
Hive
Git

Job description

We are looking for a Data Architect with strong hands-on expertise in PySpark to design, develop, optimize, and maintain scalable data processing solutions. The ideal candidate must have extensive experience writing production-grade code in PySpark and working with large-scale data processing and distributed computing environments.

Responsibilities:
  • Design, develop, and maintain scalable data pipelines using PySpark.
  • Write clean, efficient, and production-ready PySpark code for large-scale data processing.
  • Develop ETL/ELT pipelines to ingest, transform, validate, and load data from multiple sources.
  • Perform complex data transformations, aggregations, joins, filtering, and data cleansing using PySpark.
  • Optimize PySpark jobs for performance, scalability, memory utilization, and execution time.
  • Work with distributed data processing frameworks and large datasets in cloud or on-premise environments.
  • Develop reusable PySpark frameworks, libraries, and data processing components.
  • Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues.
  • Implement data validation, error handling, logging, and monitoring within data pipelines.
  • Work closely with Data Architects, Data Engineers, Data Scientists, BI teams, and business stakeholders to understand data requirements.
  • Participate in data modeling, pipeline architecture, and technical design discussions.
  • Review code and provide technical guidance and mentorship to junior and mid-level Data Engineers.
  • Ensure adherence to coding standards, data engineering best practices, security, and governance requirements.
  • Work with CI/CD processes and version control systems for deploying and managing data pipelines.
  • Contribute to technical documentation and maintain clear documentation of data pipelines and processes.
Requirements:
  • 8+ years of experience in Data Engineering or a related field.
  • Strong and mandatory hands-on experience with PySpark.
  • Proven experience writing production-grade PySpark code is mandatory.
  • Strong understanding of Spark architecture, DataFrames, Spark SQL, transformations, actions, partitioning, caching, and optimization.
  • Strong proficiency in Python with practical experience developing data engineering applications.
  • Hands-on experience building and managing large-scale ETL/ELT data pipelines.
  • Strong SQL skills with experience working on complex queries, joins, aggregations, and performance optimization.
  • Experience working with distributed data processing and big data environments.
  • Good understanding of data warehouse concepts, data modeling, and database technologies.
  • Experience with cloud data platforms such as Azure, AWS, or GCP is preferred.
  • Experience with technologies such as Databricks, Azure Data Lake, Delta Lake, Kafka, Hive, or similar big data technologies is preferred.
  • Experience with Git and CI/CD practices.
  • Strong problem-solving and analytical skills.
  • Ability to independently troubleshoot complex technical and data-related issues.
  • Good communication skills and ability to work effectively with cross-functional teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer (Pyspark, Databricks)
Senior Data Engineer (Pyspark, Databricks)

Strategic Systems International • Lahore

On-site
PKR 1,800,000 - 3,200,000
Senior PySpark Data Architect — Scalable Pipelines
Senior PySpark Data Architect — Scalable Pipelines

iSystems Ltd • Punjab

On-site
PKR 1,200,000 - 1,800,000
Senior PySpark Data Architect: Scalable Pipelines & ETL
Senior PySpark Data Architect: Scalable Pipelines & ETL

iSystems Ltd • Punjab

On-site
PKR 3,500,000 - 5,500,000
Data Engineer
Data Engineer

Zorba Consulting • Hyderabad City Taluka

On-site
INR 1,200,000 - 2,400,000
Data Architect - Microsoft Azure Data Services, DataLake, Databricks
Data Architect - Microsoft Azure Data Services, DataLake, Databricks

HireOn • Pakistan

Hybrid
PKR 3,000,000 - 5,500,000
Senior Data Engineer
Senior Data Engineer

Systems Limited • Punjab

On-site
PKR 2,400,000 - 3,600,000
Senior Manager Data
Senior Manager Data

Techsurge Private Limited • Karachi Division

On-site
PKR 3,000,000 - 6,000,000
Senior Manager Data
Senior Manager Data

TechSurge Inc • Karachi Division

On-site
PKR 4,000,000 - 6,600,000
Senior Databricks Engineer
Senior Databricks Engineer

Digifloat • Islamabad

On-site
PKR 2,600,000 - 3,400,000
Data Engineer
Data Engineer

Archisurance • Lahore

On-site
PKR 1,200,000 - 1,800,000