Lead Data Engineer

Smart Ims

Bengaluru

Hybrid

INR 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Opportunity to work with cutting-edge technologies
Collaborate with top-tier engineers

Job summary

Smart Ims in Bengaluru is seeking a skilled Data Engineer to join the Data Platform team. You will play a crucial role in building and optimizing the data infrastructure, handling petabytes of data.

The ideal candidate should have 3-5 years of data engineering experience, proficiency in Scala and Apache Spark, and familiarity with GCP and Kubernetes. Opportunities to work with cutting-edge technologies like Iceberg and GenAI are provided.

Qualifications

  • 3-5 years of hands-on experience in Data Engineering.
  • Strong proficiency in Scala and Apache Spark (Batch & Streaming).
  • Solid understanding of SQL and distributed computing concepts.

Responsibilities

  • Design, develop, and maintain robust ETL/ELT pipelines.
  • Optimize Spark jobs and SQL queries for efficiency.
  • Implement and manage data tables using Lakehouse formats.
  • Apply Medallion Architecture principles for data structuring.

Skills

Scala
Apache Spark
SQL
GCP (Google Cloud Platform)
Kubernetes
Docker
Data Engineering
Data Modeling

Education

Bachelor's or Master's degree in Computer Science, IT, Engineering or related field

Tools

Apache Iceberg
Hudi
Delta Lake
CI/CD tools
Airflow

Job description

Job description

Position: Data Engineer

Experience: 3-5 Years

Location: Bangalore (Bellandur)

Team: Data Platform & Engineering

Overview

We are looking for a skilled Data Engineer to join our Data Platform team. You will play a key role in building and optimizing our next-generation data infrastructure. Operating at the scale of Flipkart (Petabytes of data), you will design, develop, and maintain high-throughput distributed systems, bridging traditional big data engineering with modern cloud-native and AI-driven workflows.

Key Responsibilities
  • Build Scalable Pipelines: Design, develop, and maintain robust ETL/ELT pipelines using Scala and Apache Spark/Flink (Core, SQL, Streaming) to process massive datasets with low latency.
  • Performance Tuning: Optimize Spark jobs and SQL queries for efficiency, resource utilization, and speed.
  • Lakehouse Implementation: Implement and manage data tables using modern Lakehouse formats like Apache Iceberg, Hudi, or Delta Lake, ensuring efficient storage and retrieval.
  • Data Modeling: Apply Medallion Architecture principles (Bronze/Silver/Gold) to structure data effectively for downstream analytics and ML use cases.
  • Data Quality: Implement data validation checks and automated testing using frameworks (e.g., Deequ, Great Expectations) to ensure data accuracy and reliability.
  • Observability: Integrate pipelines with observability tools to monitor data health, freshness, and lineage.
  • Cloud Infrastructure: Deploy and manage workloads on GCP DataProc and Kubernetes (K8s), leveraging containerization for scalable processing.
  • Infrastructure as Code: Contribute to infrastructure automation and deployment scripts.
  • GenAI Integration: Explore and implement GenAI and Agentic workflows to automate data discovery and optimize engineering processes.
  • Agile Delivery: Work closely with architects and product teams in an Agile/Scrum environment to deliver features iteratively.
  • Code Reviews: Participate in code reviews to maintain code quality, standards, and best practices.
Required Qualifications
  • 3-5 years of hands‑on experience in Data Engineering.
  • Strong proficiency in Scala and Apache Spark (Batch & Streaming).
  • Solid understanding of SQL and distributed computing concepts.
  • Experience with GCP (DataProc, GCS, BigQuery) or equivalent cloud platforms (AWS/Azure).
  • Hands‑on experience with Kubernetes and Docker.
  • Experience with Lakehouse table formats (Iceberg, Hudi, or Delta).
  • Understanding of data warehousing and modeling concepts (Star schema, Snowflake schema).
  • Strong problem‑solving skills and ability to work independently.
  • Good communication skills to collaborate with cross‑functional teams.
  • Bachelors or Masters degree in Computer Science, Information Technology, Engineering, or a related quantitative field.
Preferred Qualifications
  • Machine Learning Background: Familiarity with ML concepts, feature engineering, or experience building data pipelines for ML models.
  • Experience with workflow orchestration tools (Airflow, Azkaban, etc.).
  • Familiarity with real‑time analytics databases (Druid, ClickHouse, HBase).
  • Experience with CI/CD pipelines for data applications.
Why Join Us?
  • Work on petabyte‑scale challenges that define the industry standard.
  • Collaborate with top‑tier engineers in a high‑growth environment.
  • Opportunity to work with cutting‑edge technologies like Iceberg, K8s, and GenAI.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Navikenz • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Lead Data Architect (AWS)
Lead Data Architect (AWS)

Anrgi Tech Private Limited • Bengaluru

On-site
INR 4,200,000 - 7,200,000
Lead Data Architect (AWS)
Lead Data Architect (AWS)

ANRGI TECH • Bengaluru

On-site
INR 4,000,000 - 6,500,000
Lead Data Engineer
Lead Data Engineer

MathCo • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Lead Data Engineer
Lead Data Engineer

SourcingXPress • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Lead Data Engineer Hyderabad, Telangana, India
Lead Data Engineer Hyderabad, Telangana, India

Cohere Health, Inc. • Hyderabad

On-site
INR 2,000,000 - 3,000,000
Lead Data Engineer
Lead Data Engineer

Weekday AI (YC W21) • Coimbatore District

On-site
INR 1,500,000 - 2,500,000
Lead Data Engineer
Lead Data Engineer

Talentrabbit • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Data Engineer
Data Engineer

Agilisium • Chennai District

On-site
INR 800,000 - 1,200,000
Lead Data Engineer (Snowflake / Databricks)Ahmedabad, Pune7–10 Years
Lead Data Engineer (Snowflake / Databricks)Ahmedabad, Pune7–10 Years

Caregence • Ahmedabad District

On-site
INR 2,500,000 - 4,200,000
Flexible Timings
5 Days Working
Healthy Environment
+4