Senior Data Engineer – Spark Streaming & Kafka

Kagool

India

On-site

INR 3,000,000 - 4,200,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Kagool in India is seeking a Senior Data Engineer with 7+ years of experience to design, develop, and support scalable data pipelines using Spark, PySpark, Kafka and Azure services.

The role emphasizes real-time streaming, Azure Databricks, ADLS Gen2, and CI/CD practices, working with cross-functional teams to deliver high-throughput analytics platforms.

Qualifications

  • 7+ years of experience in Data Engineering / Big Data.
  • Strong hands-on experience with Apache Spark and PySpark.
  • Strong hands-on experience with Apache Kafka.
  • Mandatory hands-on experience with Kafka on Azure.
  • Strong experience with Spark Structured Streaming.
  • Strong programming skills in Python.
  • Experience designing and developing real-time/streaming data pipelines.
  • Good understanding of distributed computing and Big Data architecture.

Responsibilities

  • Design, develop, and maintain scalable batch and real-time data pipelines using Apache Spark and PySpark.
  • Develop and support Kafka-based real-time streaming pipelines on Microsoft Azure.
  • Implement real-time data processing using Spark Structured Streaming.
  • Work with Kafka topics, partitions, consumer groups, offsets, retention, and message delivery mechanisms.
  • Implement Kafka producers and consumers for high-volume data ingestion and processing.
  • Integrate Apache Kafka with Azure data services and downstream data platforms.
  • Develop data transformation, cleansing, aggregation, and enrichment processes.
  • Implement checkpointing, fault tolerance, error handling, retry mechanisms, and recovery strategies for streaming applications.
  • Design and optimize Spark jobs for performance, scalability, and efficient resource utilization.
  • Work with Azure data services such as Azure Databricks, ADLS Gen2, Azure Event Hubs, Azure Data Factory, and Azure Synapse Analytics.
  • Develop SQL queries and work with relational and analytical databases.
  • Implement data quality and validation frameworks.
  • Troubleshoot production issues and perform root-cause analysis.
  • Collaborate with Data Architects, Application Teams, DevOps, and Business stakeholders.
  • Participate in technical design discussions, code reviews, and CI/CD activities.

Skills

Apache Spark
PySpark
Apache Kafka
Azure
Python
SQL
Spark Structured Streaming
Azure Databricks
Azure Data Factory
Azure Synapse Analytics

Tools

Azure Databricks
ADLS Gen2
Azure Event Hubs
Azure Data Factory
Azure Synapse Analytics

Job description

Job Summary

We are a fast-growing IT consultancy specializing in the transformation of complex Global enterprises that use SAP. We are looking for hard working individuals to help deliver for our global customer base. We embrace the opportunities of the future and work proactively to make good use of technology. As you can imagine, this means that we have a vibrant and diverse mix of skills and people making Kagool a great place to work.


We are looking for an experienced Senior Data Engineer with 7+ years of experience in designing, developing, and supporting scalable data engineering solutions.


The ideal candidate must have strong hands-on experience with Apache Spark, PySpark, Kafka, Spark Structured Streaming, Python, SQL, and Microsoft Azure. The candidate should have practical experience implementing and managing Kafka-based streaming solutions on Azure and working with real-time data processing pipelines.


Azure experience is mandatory for this role.


Key Responsibilities


  • Design, develop, and maintain scalable batch and real-time data pipelines using Apache Spark and PySpark.

  • Develop and support Kafka-based real-time streaming pipelines on Microsoft Azure.

  • Implement real-time data processing using Spark Structured Streaming.

  • Work with Kafka topics, partitions, consumer groups, offsets, retention, and message delivery mechanisms.

  • Implement Kafka producers and consumers for high-volume data ingestion and processing.

  • Integrate Apache Kafka with Azure data services and downstream data platforms.

  • Develop data transformation, cleansing, aggregation, and enrichment processes.

  • Implement checkpointing, fault tolerance, error handling, retry mechanisms, and recovery strategies for streaming applications.

  • Design and optimize Spark jobs for performance, scalability, and efficient resource utilization.

  • Work with Azure data services such as Azure Databricks, ADLS Gen2, Azure Event Hubs, Azure Data Factory, and Azure Synapse Analytics.

  • Develop SQL queries and work with relational and analytical databases.

  • Implement data quality and validation frameworks.

  • Troubleshoot production issues and perform root-cause analysis.

  • Collaborate with Data Architects, Application Teams, DevOps, and Business stakeholders.

  • Participate in technical design discussions, code reviews, and CI/CD activities.


Mandatory Technical Skills


  • 7+ years of experience in Data Engineering / Big Data.

  • Strong hands-on experience with Apache Spark and PySpark.

  • Strong hands-on experience with Apache Kafka.

  • Mandatory hands-on experience with Kafka on Azure.

  • Strong experience with Spark Structured Streaming.

  • Strong programming skills in Python.

  • Experience designing and developing real-time/streaming data pipelines.

  • Good understanding of distributed computing and Big Data architecture.


Hands-on experience with Kafka:


  • Topics and partitions

  • Producers and consumers

  • Offsets

  • Retention

  • Error handling and recovery


Strong understanding of Spark concepts including:


  • Joins

  • Caching

  • Performance tuning


Hands-on experience with Azure Databricks and/or Azure data engineering services.


Experience with Git and CI/CD practices.


Candidates must have practical experience with one or more of the following:


The candidate should have hands‑on implementation/support experience with Kafka in an Azure environment, including experience with:

Kafka deployment/integration on Azure


Kafka producers and consumers


Kafka topics and partitions


Consumer groups and offset management


Kafka-to-Spark streaming integration


Monitoring and troubleshooting Kafka workloads


Performance tuning and scalability


Security/authentication for Kafka workloads on Azure


Good to Have


  • Kafka Connect

  • Confluent Kafka / Confluent Cloud

  • Schema Registry

  • Terraform


Candidate Profile

The candidate should be comfortable working on large-scale; high-throughput streaming systems and should have experience taking data pipelines from design and development through production deployment and operational support.


Career at Kagool

A career at Kagool will give you a path towards progression and opportunities, with the current rate of growth we at Kagool have dedicated time towards individual growth, recognizing individual contributions, filling the team with a strong sense of purpose along with providing a fun, flexible and friendly work environment

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Kagool • India

On-site
INR 90,000 - 120,000
Flexible work environment
Opportunities for progression
Senior Kafka Engineer
Senior Kafka Engineer

Anveta Manpower Solutions • Dadri, Bengaluru, Mumbai

Remote
INR 4,000,000 - 7,000,000
Data Architect
Data Architect

Lancesoft • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Python & Kafka Data Engineer
Python & Kafka Data Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Senior DevOps
Senior DevOps

Luxoft • Pune District

On-site
INR 1,500,000 - 2,100,000
Senior DevOps - Cloud (Azure) - Immediate avaiable
Senior DevOps - Cloud (Azure) - Immediate avaiable

Luxoft India • Pune District

On-site
INR 3,000,000 - 6,000,000
Data Engineer
Data Engineer

Coretek Services India • Hyderabad

Hybrid
INR 2,800,000 - 4,000,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000