Data Engineering Lead

Data for Public Good

Bengaluru

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Data for Public Good in Bengaluru seeks a Data Engineering Lead to drive the architecture and execution of data onboarding across its ecosystem. You will lead a team and ensure compliance with international data policies while integrating complex datasets from various national agencies.

The ideal candidate has 5-6 years of experience in data engineering, a B.E./B.Tech in a technical field, and expertise in building resilient ETL modules.

Qualifications

  • 5 to 6 years of experience in data engineering.
  • Proven experience with real-time ETL using REST, AMQP, MQTT.

Responsibilities

  • Lead a team of data engineers and oversee resource allocation.
  • Define data protocols and strategic integration roadmaps.
  • Monitor data flows and resolve integration errors.
  • Create documentation for integration procedures.

Skills

Analytical and problem-solving skills
ETL module building
PySpark for transformations
Data processing workflows
Knowledge of geospatial data
Linux environments
Proficiency in GIT
Monitoring/logging stacks

Education

B.E./B.Tech in CSE, ECE, or related technical field

Tools

Docker
Kubernetes
QGIS

Job description

Title:Data Engineering Lead
Location:Bangalore (with occasional travel)
Experience:5 to 6 Years
Qualification:B.E./B.Tech in CSE, ECE, or related technical field

About CDPG

At CDPG, we are committed to democratizing data. Our mission is to harness its power by creating sovereign data exchange platforms that serve the Public Good. By ensuring that data exchange is ethical, secure, and privacy-centric, we strive to make the benefits of data accessible to all, promoting regional inclusivity and informed decision-making.

Scope of Work

As the Data Engineering Lead, you will own the end-to-end strategy, architecture, and execution of data onboarding across the CDPG ecosystem. You will lead a high-performing team of engineers to integrate complex datasets from national agencies, government departments, and system integrators. Your mission is to manage the full data lifecycle, ensuring all integrations are scalable, secure, and fully compliant with international data sovereignty policies.

Responsibilities
  • Lead a team of data engineers, oversee resource allocation, and drive technical upskilling in AI-driven development and automation.
  • Act as the primary technical liaison for high-level government officials and agency heads to define data protocols and strategic integration roadmaps.
  • Evaluate APIs and datasets to design robust data models (JSON-LD, GeoJSON). Architect ETL modules that are containerized and ready for deployment.
  • Oversee the development of Python-based ETL modules for data ingestion via REST APIs and streaming protocols (AMQP, MQTT).
  • Own the integration delivery lifecycle using Agile methodologies, ensuring high-quality, on-time delivery of data flows.
  • Monitor data flows to identify and resolve integration errors, ensuring high availability and performance.
  • Create and maintain documentation for integration procedures, data maps, and technical specifications.
  • Enforce software development best practices, ensuring rigorous unit/functional testing and maintaining the integrity of the Data Catalogue.
Technical Skills
  • Strong analytical and problem-solving skills to troubleshoot complex data issues. Ability to evaluate data availability, frequency, and systemic issues (e.g., stream repetitions or schema drift) within data systems.
  • Proven experience building ETL modules for real-time ingestion using REST, AMQP, and MQTT protocols.
  • Expertise in building resilient data processing workflows using PySpark for heavy lift transformations and Airflow, NiFi for orchestration.
  • Strong knowledge of modern data warehouse architectures: Medallion with Bronze, Silver, Gold refinement layers, Data Mesh with decentralized, domain-owned data products.
  • Ability to leverage AI technologies and LLMs within data integration frameworks to streamline and accelerate the integration lifecycle.
  • Good command of geospatial data (vector/raster) and GIS tools such as QGIS or similar for analyzing and transforming spatial datasets is desirable.
  • Proficient in Linux environments and GIT version control. Familiarity with Docker and Kubernetes deployment orchestration is desirable.
  • Experience with monitoring/logging stacks Prometheus, Grafana, and Logstash to ensure system health and reliability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Cynosure Corporate Solutions • Chennai District

On-site
INR 400,000 - 600,000
GCP Data Engineer
GCP Data Engineer

CustomerLabs • Chennai

On-site
INR 800,000 - 1,200,000
Lead Data Engineer
Lead Data Engineer

Ada Digital Analytics • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Data Engineer
Data Engineer

Deservely Technologies Pvt Ltd • Hyderabad

On-site
INR 1,500,000 - 2,300,000
Delivery Lead/Senior Data Engineer 3
Delivery Lead/Senior Data Engineer 3

1203 Barclays Global Serv. Cent • Pune District

On-site
INR 2,500,000 - 3,500,000
Data Engineer
Data Engineer

Capri Global Capital (CGCL) • Gurugram District, Dadri

Hybrid
INR 900,000 - 1,300,000
Data Engineer - Gurugram
Data Engineer - Gurugram

Yeah! Global • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000
Data Engineer
Data Engineer

1203 Barclays Global Serv. Cent • Pune District

On-site
INR 1,000,000 - 1,500,000
Data Engineer (Contract)
Data Engineer (Contract)

NexTurn Inc. • Hyderabad

On-site
INR 1,200,000 - 1,800,000