Designation: Lead Data Engineer (Databricks & PySpark)
Experience: 8 to 14 years
Location: Mumbai
Work Mode: Hybrid
About the Role: We are seeking a highly skilled and experienced Lead Data Engineer to design, build, and operate our next-generation data platform. In this role, you will champion Legacy Code Modernization efforts, migrating traditional ETL processes into modern, scalable ELT patterns on Databricks. As a technical leader, you will manage a cross‑functional team, oversee end-to-end delivery, and collaborate with Infrastructure, Applications, and Cyber Security teams to drive data engineering excellence across the organization.
Key Responsibilities:
- Design and Build: Develop reliable, scalable end-to-end data workflows from ingestion through transformation to consumption on the Databricks platform.
- Operations & Monitoring: Implement robust error handling, alerting mechanisms, and monitoring to ensure pipeline performance and uptime.
- Performance Tuning: Optimize Spark job design and cluster configurations to maximize throughput and minimize cloud infrastructure costs.
- Orchestration: Manage complex multi‑stage data workflows using Databricks Jobs and modern orchestration patterns.
- Legacy Code Modernization
- Code Refactoring: Assess existing, traditional codebases and refactor legacy ETL workflows into highly efficient PySpark pipelines.
- ELT Migration: Transition legacy data structures to modern ELT lakehouse patterns on Databricks.
- Risk Mitigation: Maintain backward compatibility and data integrity during migrations, creating clear playbooks to minimize business disruption.
- Data Engineering Excellence & Leadership
- Governance & Quality: Implement data quality verification frameworks and validation checks to protect data integrity.
- Delta Lake Architecture: Design and optimize Delta Lake tables utilizing advanced storage features, ACID transactions, and schema evolution.
- Team Leadership: Manage, mentor, and foster growth for junior and mid‑level data engineers, driving knowledge‑sharing initiatives across Team.
- Delivery Management: Own project delivery by participating in agile ceremonies, sprint planning, estimation, and release management.
Job Requirements Experience & Background:
- Overall Experience: 8 to 14 years of professional experience in data engineering or related backend fields.
- Databricks Focus: 4 to 8 years of intensive, hands‑on experience building scale production workflows on Databricks.
- Leadership: 2 to 3 years of direct experience handling engineering teams, managing stakeholders, and providing production support.
Essential Technical Skills:
- PySpark
- Advanced Python programming capabilities tailored for data engineering and automation workflows.
- SQL: Deep proficiency in writing complex SQL queries, analytical functions, and data transformations.
- Delta Lake: Solid understanding of Delta Lake optimization strategies, transactional layouts, and data modeling (dimensional, data vault, or lakehouse).
- Workspace AI Agent: Familiarity with Databricks Workspace AI Agent features and integration workflows.
Desirable Technical Skills (Nice to Have):
- Cloud Architecture
- DevOps/DataOps: Streaming.
Mandatory Certifications:
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional Preferred
- Databricks Certified Associate Developer for Apache Spark
- Cloud platform data certifications (e.g., Azure Data Engineer Associate, AWS Certified Data Analytics, GCP Professional Data Engineer.)