Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD

Gurugram District

On-site

INR 4,000,000 - 7,500,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Srijan Technologies PVT LTD in Gurugram seeks a Lead Data Engineer to architect and own scalable data pipelines across Databricks, Snowflake, AWS, and Azure, guiding a skilled team toward production-ready solutions.

You will implement CDC and real‑time processing, build REST APIs, collaborate with ML teams on data readiness, and strengthen governance, security, and CI/CD with GitLab, Docker, and Kubernetes.

Qualifications

  • 5+ years of experience in Data Engineering.

Responsibilities

  • Lead development of scalable ETL/ELT pipelines and data models for large retail data.
  • Design and optimize multi-platform data processing across Databricks, Snowflake, AWS, and Azure.
  • Build data pipelines with Airflow and Airbyte for reliable data movement.
  • Implement CDC, batch, and real-time processing with scalable architectures.
  • Develop and manage REST APIs and external data ingestion from APIs.
  • Collaborate with ML teams to prepare data for model deployment.
  • Own CI/CD pipelines with Git, Docker, and Kubernetes; ensure governance and security.

Skills

Python
SQL
ETL/ELT
Airflow
Airbyte
Databricks
Snowflake
AWS
Azure
MLOps
CI/CD
Docker
Kubernetes
APIs REST
FastAPI/Flask
Delta Lake
CDC
Data Governance

Tools

Airflow
Airbyte
Docker
Kubernetes

Job description

Lead Data Engineer

Overview

We are looking for a Lead Data Engineer who combines hands‑on multi‑platform expertise with strong leadership in data architecture, pipelines, and CI/CD. This role requires a versatile engineer with deep technical skills across modern data platforms (such as Databricks, Snowflake, AWS, and Azure), an understanding of MLOps/DevOps practices, and the ability to guide a high‑performing team in building scalable, production‑ready data solutions. You will not be limited to a single platform but will leverage a diverse toolkit to solve complex data challenges.

Key Responsibilities
  • Pipeline & Architecture: Lead hands‑on development of scalable ETL/ELT pipelines, data models, and integration frameworks to process high‑volume (billions of records) structured and unstructured retail data.
  • Multi‑Platform Engineering: Design, develop, and optimize data processing applications across multiple platforms, including Databricks (Spark/Delta Lake), Snowflake, AWS, or Azure.
  • Data Integration & Orchestration: Build and manage robust data pipelines using Apache Airflow for orchestration and Airbyte for seamless data integration and movement.
  • Data Processing: Architect and implement robust solutions for Change Data Capture (CDC), large‑scale batch processing, and low‑latency real‑time/streaming data processing.
  • API Management: Work extensively with external APIs for data ingestion, as well as design, create, and manage internal REST APIs to serve data to downstream applications and users.
  • AI‑Augmented Deliverables: Actively leverage AI assistants to conceptualize, design, and accelerate the development of data pipelines and everyday engineering tasks.
  • DevOps & CI/CD: Own and evolve CI/CD pipelines (Git workflows, automated testing, release cycles, secrets management, documentation). Guide DevOps‑oriented deployments utilizing Dockerized applications, Kubernetes orchestration, and monitoring/logging tools (Splunk, Datadog, Dynatrace).
  • MLOps Alignment: Collaborate with Data Scientists on data readiness for ML projects and ensure alignment with ML lifecycle stages (data prep, feature engineering, model deployment).
  • Governance & Leadership: Establish and enforce best practices in data governance, data quality, metadata, and security. Mentor team members through peer reviews, knowledge sharing, and technical leadership.
  • Innovation: Stay ahead of industry trends in MLOps, observability, and GenAI, introducing relevant tools and practices.
Required Skills & Experience
  • Experience: 5+ years of experience in Data Engineering.
  • Data Lakes & Warehouses: Mandatory expertise in designing, building, and managing large‑scale Data Warehouses and Data Lakes from the ground up.
  • Data Processing Paradigms: Extensive, hands‑on experience working with Change Data Capture (CDC) mechanisms, complex batch processing, and real‑time/streaming data processing.
  • Platform Expertise: Proven expertise in more than one major cloud data platform/ecosystem (e.g., Databricks, Snowflake, AWS Analytics, Azure Data Engineering).
  • SQL Mastery: Advanced proficiency in writing, optimizing, and debugging complex SQL queries for large‑scale data processing and analytics.
  • Programming: Strong programming skills in Python (async, threading, decorators, advanced I/O).
  • APIs: Strong proficiency in interacting with third‑party APIs and hands‑on experience creating and managing REST APIs (using frameworks like FastAPI, Flask, or similar).
  • Tooling: Deep hands‑on experience with workflow orchestration (Apache Airflow) and data integration platforms (Airbyte).
  • AI‑Assisted Engineering: Mandatory capability to use AI coding assistants and tools to design pipelines, write code, and enhance day‑to‑day productivity.
  • Data Architecture: Experience with data modeling (e.g., Delta Lake or Snowflake architecture) and scalable ETL/ELT design.
  • DevOps/CI/CD: Hands‑on experience with Git‑based CI/CD (GitLab preferred) and a working knowledge of Docker & Kubernetes for deployment and scaling.
  • MLOps: Understanding of MLOps concepts including data preparation, model lifecycle, registries, and monitoring.
  • Soft Skills: Strong problem‑solving skills with the ability to design for scale and performance, coupled with excellent collaboration, communication, and leadership skills.
Good to Have
  • Customer Data Platform (CDP): Experience working with, building, or implementing CDPs to unify customer data across systems.
  • Experience in the retail domain or other large‑scale data‑heavy environments.
  • Familiarity with streaming frameworks (Kafka, Spark Streaming, etc.).
  • Agentic Pipeline Development: Experience or strong interest in building agentic pipelines using LLMs for dynamic data orchestration and automation.
  • Knowledge of model observability tools and ML deployment pipelines.
  • Exposure to GenAI concepts (vector embeddings, vector databases, RAG).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Technical Lead - Data Engineer
Technical Lead - Data Engineer

Srijan: Now Material • Gurugram District

On-site
INR 2,500,000 - 4,000,000
Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer - Lead
Data Engineer - Lead

Iris Software • Dadri

On-site
INR 1,500,000 - 2,500,000
Principal Engineer - Data Engineer
Principal Engineer - Data Engineer

Staples India • Chennai District

On-site
INR 3,500,000 - 7,000,000
Senior Data Engineer (AI/ML)
Senior Data Engineer (AI/ML)

Neolatika • Maharashtra

On-site
INR 2,000,000 - 3,600,000
Lead Data Engineer
Lead Data Engineer

Talentrabbit • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000
Senior Data Engineer
Senior Data Engineer

KSB • Pune District

On-site
INR 800,000 - 1,200,000
Data Science & Engineering Lead
Data Science & Engineering Lead

Particle41 • Pune District

On-site
INR 8,555,000 - 12,358,000