Lead Data Engineer

Relanto

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Relanto is seeking a Lead Data Engineer with deep hands-on experience in real-time streaming, event processing, CDC, and modern data engineering. You will design, build, and maintain high-performance data pipelines using Apache Flink, Kafka, Debezium, and ClickHouse, plus orchestration with Airflow.

The role covers end-to-end data flow from source systems through Kafka and Flink processing to analytics storage, dashboards, and downstream integrations.

Qualifications

  • 6–8 years of experience in Data Engineering.
  • Strong hands-on experience with Apache Flink (mandatory).
  • Strong hands-on experience with Apache Kafka.
  • Hands-on experience with Debezium and CDC.
  • Strong programming in Java or Scala.
  • Experience with ClickHouse or analytical databases.
  • Experience with Apache Airflow or similar orchestration.
  • Distributed systems and real-time data processing knowledge.

Responsibilities

  • Design, develop, and maintain real-time data pipelines using Flink and Kafka.
  • Build production-grade streaming applications for high-volume, low-latency workloads.
  • Implement data transformation, enrichment, filtering, and aggregation.
  • Develop and maintain Flink jobs for real-time processing.
  • Monitor latency, throughput, failures, and resource use; troubleshoot issues.

Skills

Apache Flink
Apache Kafka
Debezium
CDC
Java
Scala
Python
ClickHouse
Apache Airflow
Kubernetes
Docker
Cloud platforms
CI/CD
REST APIs
Data pipelines

Tools

Kafka tooling
Debezium connectors
Airflow DAGs

Job description

We are looking for a Lead Data Engineer with strong hands-on experience in real-time data streaming, event processing, CDC, data integration, and modern data engineering.

The Lead Data Engineer will be responsible for designing, developing, and maintaining highperformance data pipelines using Apache Flink, Apache Kafka, Debezium, CDC, ClickHouse, and Apache Airflow.

This is a hands-on engineering role requiring strong practical experience in building productiongrade streaming applications, developing data pipelines, troubleshooting distributed systems, and optimizing data processing workloads.

The ideal candidate should be comfortable working across the full data pipeline—from source systems and CDC ingestion through Kafka and Flink processing to analytical storage, APIs, dashboards, and downstream integrations.

Key Responsibilities

  • Design, develop, and maintain real-time data pipelines using Apache Flink and Apache Kafka.
  • Develop production-grade streaming applications for high-volume and low-latency workloads.
  • Implement data transformation, filtering, enrichment, aggregation, and event processing.
  • Build reliable event-processing pipelines with appropriate error handling and recovery mechanisms.
  • Consume and publish events across Kafka topics. Implement appropriate partitioning, consumer groups, offsets, and delivery mechanisms.
  • Troubleshoot streaming pipeline failures and performance issues.
  • Develop and maintain Apache Flink jobs forreal-time data processing.
  • Implement:
    • Stream transformations
    • Filtering
    • Mapping
    • Aggregations
    • Joins
    • Windows
    • Event-time processing
    • Watermarks
    • State management
  • Implement Flink checkpointing and recovery mechanisms.
  • Optimize Flink jobs for performance, scalability, and resource utilization.
  • Monitor Flink jobs for latency, throughput, failures, and resource consumption.
  • Troubleshoot state, checkpointing, backpressure, and processing issues.
Kafka
  • Develop Kafka-based ingestion and streaming pipelines.
  • Create and manage Kafka topics and event streams.
  • Workwith partitions, offsets,consumer groups, replication, and retention.
  • Implement reliableproducer and consumer applications.
  • Handle message ordering, retries,duplicate events, and replay scenarios.
  • Monitor Kafka performance and troubleshoot consumerlag and throughput issues.
  • Work with Kafka schemasand serialization formats.
  • BuildCDC-based ingestion pipelines using Debezium.
  • Configure and maintain Debeziumconnectors.
  • Capture source-system inserts, updates,and deletes.
  • Publish CDC events into Kafka.
  • Handle initial snapshots and incremental CDC processing.
  • Manage schemaevolution and changesin source systems.
  • Implement data reconciliation and consistency checks.
  • Troubleshoot CDC failures and source-to-target data issues.
Data Orchestration
  • Develop and maintain data workflows using Apache Airflow orequivalent orchestration frameworks.
  • Data validation
  • Flink job execution
  • Data transformation
  • Downstream integrations
  • Implement workflowdependencies, scheduling, retries,backfills, SLAs, and alerting.
  • Integrate Airflow workflows with Kafka, Flink, Debezium, ClickHouse, APIs, and cloud services.
  • Monitor workflow execution and troubleshoot failures.
  • Develop reusable operators, sensors,and workflow components where appropriate.
  • Useevent-driven triggers where real-time workflows require them.
ClickHouse & Analytical Data
  • Integrate streaming data pipelines with ClickHouse.
  • Design efficient analytical data models.
  • Develop and optimize SQL queries.
  • Implement appropriate partitioning, sorting, indexing,and retention strategies.
  • Optimize data ingestion and query performance.
  • Support downstream analytical use cases,dashboards, and reporting requirements.
Data Quality& Reliability
  • Implement data validation and quality checksthroughout the pipeline.
  • Buildreconciliation mechanisms betweensource and target systems.
  • Monitor data freshness, completeness, accuracy, and consistency.
  • Implement error handling, retry,replay, and recoverymechanisms.
  • Establish loggingand observability for critical pipelines.
  • Support incident investigation and root-cause analysis.
Integration & APIs
  • Integrate streaming and analytical data with APIs, endpoints, dashboards, and downstream applications.
  • Develop data interfaces and integration components.
  • Workwith application teams to define data contractsand integration requirements.
  • Support future integrations and additional data consumers.
  • Follow modernsoftware engineering practices around:
    • Git
    • Code reviews
    • Unit testing
    • Integration testing
    • CI/CD
    • Logging
    • Monitoring
    • Documentation
  • Develop reusable and maintainable data engineering components.
  • Participate in technical designdiscussions and architecture reviews.
  • Mentor other Data Engineers and contribute to engineering standards.
Required Skills& Experience
  • 6–8 years ofexperience in Data Engineering.
  • Strong hands-onexperience with Apache Flink — mandatory/core requirement.
  • Strong hands-on experience with Apache Kafka.
  • Hands-on experience with Debeziumand CDC.
  • Strong programming experience in Java or Scala.
  • Python experience is an advantage.
  • Experience with analytical databases; ClickHouse is highly preferred.
  • Hands-on experience with Apache Airflowor another data orchestration framework.
  • Strong understanding of distributed systemsand real-time data processing.
  • Experience developing and supporting production-grade streaming pipelines.
  • Strong understanding of Kafka:
    • Topics
    • Partitions
    • Offsets
    • Consumer groups
    • Replication
    • Retention
  • Experience with data transformation, enrichment, filtering, aggregation, and event processing.
  • Experience troubleshooting performance and reliability issues.
  • Familiarity with cloud platforms and containerized environments.
Preferred Skills
  • Advanced experience with Apache Flink, including:
    • State management
    • Checkpoints
    • Savepoints
    • Watermarks
    • Event time
    • Windows
    • Backpressure
    • State backends
  • Experience with Apache Airflow, Dagster,Prefect, or Apache NiFi.
  • Experience with Kafka Schema Registry.
  • Experience with Avro, Protobuf, or JSON.
  • Experience with Kubernetes.
  • Experience with AWS / Azure / GCP.
  • Experience with CI/CD pipelines.
  • Experience with Docker and containerized applications.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with data observability and monitoring tools.
  • Experience buildinghigh-volume, low-latency real-time data platforms.
  • Experience with REST APIs and systemintegrations.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Sourcebae • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Data Architect
Data Architect

Relanto • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Sr. Data Engineer
Sr. Data Engineer

AMISEQ • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Data Engineer (Flink & Iceberg)
Data Engineer (Flink & Iceberg)

Navikenz • Bengaluru

On-site
INR 900,000 - 1,600,000
Data Engineer - Real Time Streaming
Data Engineer - Real Time Streaming

GSSTech Group • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Engineer – Real-Time Streaming (Flink / Java / Kafka)
Data Engineer – Real-Time Streaming (Flink / Java / Kafka)

D4 Insight • Bengaluru

On-site
INR 900,000 - 1,300,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

Intercontinental Exchange Holdings, Inc. • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE Clear Europe Limited • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,500,000 - 2,000,000