Data Engineering Full-time, Onsite Bangalore / Chennai / Kochi (Work from Office)
Exp: 4-6 years experience
About the Role
We are looking for a Senior Data Engineer to design, build, and operate the data pipelines and lakehouse platform that power analytics and data products across the business. You will own end-to-end ingestion, transformation, and modeling on Databricks and Microsoft Fabric - turning raw, heterogeneous source data into reliable, well-governed, analytics-ready datasets.
This is a hands‑on senior individual‑contributor role. Beyond writing production-grade pipelines, you will set engineering standards, review designs, and mentor mid-level engineers on the team. The role is industry-agnostic — you will work across a range of data domains and source systems.
What You’ll Do
- Build & operate pipelines - Design and maintain scalable batch and streaming data pipelines (ELT/ETL) on Databricks, ingesting from databases, files, APIs, and event streams.
- Lakehouse & modeling - Architect and implement a medallion (bronze/silver/gold) lakehouse on Delta Lake; design dimensional and analytics-ready data models for downstream consumption.
- Microsoft Fabric delivery - Deliver data into and across Microsoft Fabric (Lakehouse, Warehouse, OneLake, Pipelines/Dataflows) and enable BI and semantic-model consumption.
- Orchestration & reliability - Orchestrate workflows with Airflow and/or Azure Data Factory; build in data‑quality checks, monitoring, alerting, and SLAs.
- Governance - Implement cataloging, lineage, and access controls using Unity Catalog and Microsoft Purview.
- Performance & cost - Tune Spark jobs, storage layout, and cluster configuration for performance and cost efficiency.
- Standards & mentorship - Define engineering best practices, conduct design and code reviews, and mentor mid-level data engineers.
What We’re Looking For
- 4–6 years of professional data engineering experience building and operating production data pipelines.
- Strong hands‑on experience with Databricks - Spark (PySpark/Spark SQL), Delta Lake, notebooks/jobs, and cluster management.
- Proven experience designing and building data pipelines (batch and streaming) at scale, with a solid grasp of ELT/ETL patterns.
- Experience with Microsoft Fabric - Lakehouse, Warehouse, OneLake, and Fabric Pipelines/Dataflows.
- Strong data modeling skills - medallion/lakehouse architecture and dimensional modeling.
- Proficiency in Python and advanced SQL.
- Experience with workflow orchestration (Airflow and/or Azure Data Factory).
- Experience with data governance tooling (Unity Catalog and/or Microsoft Purview).
- Experience with streaming ingestion (Kafka / Event Hubs and Spark Structured Streaming).
- Ability to work independently, set technical direction, and mentor others.
Nice to Have
- AWS data experience - S3, Glue, Redshift, EMR, Athena, or Lake Formation (a strong plus).
- Infrastructure-as-code and CI/CD for data (Terraform, Git-based deployment, automated testing).
- Experience with dbt for transformation and testing.
- Familiarity with Power BI semantic models and BI enablement.
- Databricks or Microsoft (Azure) certifications.
Microsoft / Fabric:
Microsoft Fabric (Lakehouse, Warehouse, OneLake, Pipelines, Dataflows Gen2), Azure Data Factory, Azure Data Lake Storage, Microsoft Purview, Power BI.
Languages & Tooling:
Python, SQL, Git, CI/CD, Terraform, dbt.
AWS (plus):
S3, AWS Glue, Amazon Redshift, EMR, Athena, Lake Formation.
Why Join Us
You’ll take real ownership of a modern lakehouse platform, work across a broad set of data domains, and shape how the team builds data. We value skills and impact over formal credentials — there is no degree or certification requirement for this role.