Lead Data Engineer

ORMAE

Pune District

On-site

INR 1,500,000 - 2,100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ORMAE in Pune, India is seeking a senior data engineer to design and implement batch, near-real-time and streaming data pipelines on Azure. You will build Bronze, Silver, and Gold data layers using Medallion Architecture and deploy PySpark workloads on Kubernetes or AKS.

You will optimize Spark jobs, implement governance, observability, and CI/CD pipelines for Spark apps and Docker images, collaborating with data engineers to uphold quality and scalability.

Qualifications

  • 6+ years of hands-on data engineering experience.
  • Strong experience with Microsoft Azure data platforms.
  • Advanced Python, PySpark, and SQL skills.
  • Strong hands-on experience with Spark, Kubernetes and AKS, Docker, Azure Data Lake Storage Gen2, Azure Event Hubs, Azure DevOps and Git.
  • Experience deploying and operating Spark applications on Kubernetes.
  • Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast joins, shuffle optimization, and skew handling.
  • Experience with batch, streaming, ETL/ELT, and event-driven processing patterns.
  • Experience implementing Medallion Architecture.
  • Experience with REST APIs, SFTP, JSON, CSV, Parquet, Delta Lake, and relational databases.
  • Experience with CDC, incremental processing, schema enforcement, and schema evolution.
  • Strong understanding of data modelling, schema design, partitioning, and storage optimization.
  • Experience with Kubernetes Jobs, Cron Jobs, Config Maps, Secrets, resource limits, node pools, and autoscaling.
  • Experience implementing pipeline observability, data validation, monitoring, alerting, and error recovery.

Responsibilities

  • Design and build batch, near-real-time, and streaming data pipelines on Azure.
  • Develop Bronze, Silver, and Gold data layers using Medallion Architecture.
  • Build and deploy containerized PySpark workloads on Kubernetes or AKS.
  • Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.
  • Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.
  • Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.
  • Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.
  • Optimize Spark jobs, partitioning, shuffles, joins, file sizes, and query performance.
  • Implement monitoring, logging, alerting, audit controls, and data-quality checks.
  • Build reusable Python, PySpark, and SQL components.
  • Create CI/CD pipelines for Spark applications, Docker images, and Kubernetes deployments.
  • Review technical designs and support data engineers with implementation standards.

Skills

Python
PySpark
SQL
Azure data platforms
Kubernetes
AKS
Docker
Spark tuning
ETL/ELT
Data modeling

Tools

Azure Data Lake Storage Gen2
Azure Event Hubs
Delta Lake
Git & Azure DevOps
Docker
Kubernetes
Spark on Kubernetes

Job description

Key Responsibilities


  • Design and build batch, near-real-time, and streaming data pipelines on Azure.

  • Develop Bronze, Silver, and Gold data layers using Medallion Architecture.

  • Build and deploy containerized PySpark workloads on Kubernetes or AKS.

  • Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.

  • Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.

  • Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.

  • Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.

  • Optimize Spark jobs, partitioning, shuffles, joins, file sizes, and query performance.

  • Implement monitoring, logging, alerting, audit controls, and data-quality checks.

  • Build reusable Python, PySpark, and SQL components.

  • Create CI/CD pipelines for Spark applications, Docker images, and Kubernetes deployments.

  • Review technical designs and support data engineers with implementation standards.


Required Skills


  • 6+ years of hands-on data engineering experience.

  • Strong experience with Microsoft Azure data platforms.

  • Advanced Python, PySpark, and SQL skills.

  • Strong hands-on experience with: Apache Spark, Kubernetes and AKS, Docker, Azure Data Lake Storage Gen2, Azure Event Hubs, Azure DevOps and Git

  • Experience deploying and operating Spark applications on Kubernetes.

  • Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast joins, shuffle optimization, and skew handling.

  • Experience with batch, streaming, ETL, ELT, and event-driven processing patterns.

  • Experience implementing Medallion Architecture.

  • Experience with REST APIs, SFTP, JSON, CSV, Parquet, Delta Lake, and relational databases.

  • Experience with CDC, incremental processing, schema enforcement, and schema evolution.

  • Strong understanding of data modelling, schema design, partitioning, and storage optimization.

  • Experience with Kubernetes Jobs, Cron Jobs, Config Maps, Secrets, resource limits, node pools, and autoscaling.

  • Experience implementing pipeline observability, data validation, monitoring, alerting, and error recovery.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Azure Senior Data Engineer - S
Azure Senior Data Engineer - S

Tata Consultancy Services • Chennai District, Bengaluru, Kolkata District

On-site
INR 2,500,000 - 4,200,000
Senior Data Engineer
Senior Data Engineer

SOTI Inc • Ernakulam

On-site
INR 1,500,000 - 2,000,000
Azure Data Engineer
Azure Data Engineer

Epergne Solutions • Chennai District

On-site
INR 1,800,000 - 2,400,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Senior Data Engineer (Jaipur, India)
Senior Data Engineer (Jaipur, India)

Thoughtswinsystems • Jaipur

On-site
INR 1,200,000 - 2,000,000
Sr. Data Engineer
Sr. Data Engineer

iLink Digital • Pune District

On-site
INR 1,800,000 - 3,200,000
Sr. Data Engineer
Sr. Data Engineer

iLink Digital • Pune District

On-site
INR 1,500,000 - 2,000,000
Senior Data Engineer
Senior Data Engineer

Toppan Merril • Chennai District

On-site
INR 800,000 - 1,200,000