Senior Manager

Ex

Pune District

Hybrid

INR 1,500,000 - 2,100,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ex in Pune region is seeking a Databricks Engineer with 9–12 years of experience to design and run scalable data lakehouse solutions. You will build and manage Databricks workspaces, clusters, and jobs, implement Delta Live Tables, and govern data using Unity Catalog across cloud environments.

You will optimize PySpark pipelines, implement ML workflows with MLflow, and enable Generative AI features while ensuring cost efficiency and robust data quality.

Qualifications

  • Bachelor’s or master’s degree in computer science, information technology, data engineering, or related field.
  • 4+ years of total experience in data engineering or software engineering.
  • 3+ years of hands‑on experience with the Databricks platform in production environments.
  • Strong background in big data engineering, cloud data platforms, and distributed computing.
  • Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and Repos.
  • Proficiency with Unity Catalog — metastore setup, catalog/schema/table management, access controls, and data lineage.
  • Hands-on experience with Delta Live Tables (DLT) — pipeline development, expectations, and monitoring.
  • Experience with Spark performance tuning — AQE, query plans, partitioning, and caching.
  • Knowledge of cloud platforms and networking for Databricks (VNet/VPC, private endpoints, firewalls).

Responsibilities

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev, test, and prod.
  • Configure and manage Unity Catalog for data governance, access control, metadata, and lineage.
  • Optimize cluster configurations, including instance types, auto-scaling, spot nodes, and compute pools.
  • Implement workspace best practices, secret management, and integration with cloud key vaults.
  • Create, schedule, and manage Databricks Jobs, Workflows, and multi-task orchestration.
  • Design Delta Lake tables with partitioning, Z-Ordering, OPTIMIZE, VACUUM, and compaction.
  • Build Medallion Architecture (Bronze–Gold) for scalable data lakehouse solutions.
  • Develop Delta Live Tables pipelines with data quality expectations for ETL/ELT.
  • Manage schema evolution, Time Travel, CDF, and versioning for incremental processing.
  • Design lakehouse architectures integrating Delta Lake with ADLS, Kafka, Event Hubs, and Kinesis.
  • Develop scalable batch and real-time pipelines using PySpark, Spark SQL, and Structured Streaming.
  • Build streaming ingestion pipelines from Kafka, Event Hubs, and other sources; optimize PySpark with AQE and Photon.
  • Create reusable transformation frameworks and templates to boost productivity and standardization.
  • Implement robust error handling, logging, monitoring, and DLQ patterns for production pipelines.
  • Set up MLflow experiments, model registry, and model lifecycle management; support ML workloads and GPUs.
  • Develop feature pipelines with Databricks Feature Store; enable Generative AI capabilities including RAG.
  • MLOps practices: model versioning, deployment, A/B testing, and Databricks Model Serving.
  • Integrate Databricks with ADLS and other cloud-native services; maintain CI/CD for notebooks, jobs, and workflows.
  • Automate infrastructure deployment using DABs, Terraform, and IaC; manage data ingestion via Auto Loader and dbt.
  • Monitor pipelines, cluster usage, performance, and cloud costs; enforce data quality via Delta Live Tables and GE.
  • Implement row-level security, column masking, and governance policies; ensure end-to-end data lineage.

Skills

Databricks
PySpark
Delta Lake
Unity Catalog
Delta Live Tables
Apache Spark
Python
Cloud platforms (Azure/AWS/GCP)

Education

Bachelor's or Master's in Computer Science / IT / Data Engineering

Tools

Terraform
Fivetran
dbt
Airbyte

Job description

Role: Databricks Engineer

Experience: 9-12 Years

Location: ALL EXL Locations

Work Mode: Hybrid

Key Role and Responsibilities:

Design, build, and maintain Databricks workspaces, clusters, and compute pools across development, testing, and production environments.

Configure and manage Unity Catalog for data governance, fine-grained access control, permissions, metadata management, and data lineage.

Optimize Databricks cluster configurations, including instance types, auto-scaling, spot/preemptible nodes, and compute pools to improve performance and reduce costs.

Implement workspace best practices, including folder structures, access controls, secret management using Databricks Secrets, Azure Key Vault, or AWS Secrets Manager.

Create, schedule, and manage Databricks Jobs, Workflows, and multi-task job orchestration with dependency management.

Design and implement Delta Lake tables using partitioning, Z-Ordering, OPTIMIZE, VACUUM, and file compaction techniques.

Build and maintain Medallion Architecture (Bronze, Silver, and Gold layers) for scalable and governed data lakehouse solutions.

Develop Delta Live Tables (DLT) pipelines with built-in data quality expectations for reliable ETL/ELT processing.

Manage schema evolution, table versioning, Time Travel, and Change Data Feed (CDF) to support incremental data processing.

Design and implement lakehouse architectures integrating Delta Lake with cloud storage and external systems such as Azure Data Lake Storage (ADLS), Kafka, Event Hubs, and Kinesis.

Develop scalable batch and real-time data pipelines using PySpark, Spark SQL, Structured Streaming, and Delta Lake.

Build streaming ingestion pipelines from Kafka, Azure Event Hubs, and other streaming platforms into Delta tables.
Optimize PySpark applications using broadcast joins, Adaptive Query Execution (AQE), dynamic partition pruning, caching, and Photon Engine.

Develop reusable transformation frameworks, utility libraries, and pipeline templates to improve engineering productivity and standardization.

Implement robust error handling, retry mechanisms, logging, monitoring, and dead-letter queue (DLQ) patterns for production-grade pipelines.

Set up and manage MLflow experiment tracking, model registry, and model lifecycle management.

Support machine learning workloads by enabling scalable model training, inference, and GPU-based compute environments.

Develop feature engineering pipelines using Databricks Feature Store to create reusable and versioned machine learning features.

Enable Generative AI solutions, including Retrieval-Augmented Generation (RAG), vector search, LLM fine-tuning, and Mosaic AI capabilities.

Implement MLOps best practices, including model versioning, model deployment, A/B testing, and Databricks Model Serving.

Integrate Databricks with Azure Data Lake Storage (ADLS) and other cloud-native services.

Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or GitLab CI for Databricks notebooks, jobs, and workflows.

Automate Databricks infrastructure deployment using Databricks Asset Bundles (DABs), Terraform, and Infrastructure-as-Code (IaC) practices.

Build and manage data ingestion frameworks using Auto Loader, COPY INTO, and third party integration tools such as Fivetran, dbt, and Airbyte.

Monitor pipeline execution, cluster utilization, system performance, and cloud costs using Databricks system tables and cloud monitoring tools.

Implement row-level security, column-level masking, dynamic views, and governance policies using Unity Catalog.

Enforce data quality through Delta Live Tables expectations and Great Expectations frameworks.

Perform query optimization, execution plan analysis, caching strategies, and performance tuning to improve workload efficiency.

Maintain enterprise data cataloging, metadata management, and end-to-end data lineage.

Prepare technical documentation, architecture diagrams, operational runbooks, and standard operating procedures for Databricks platform and data engineering solutions.

Qualifications for Candidates

Bachelor’s or master’s degree in computer science, Information Technology, Data Engineering, or related field.

4+ years of total experience in data engineering or software engineering.

3+ years of dedicated hands-on experience with the Databricks platform in production environments.

Strong background in big data engineering, cloud data platforms, and distributed computing.

Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and Repos.

Proficiency with Unity Catalog — metastore setup, catalog/schema/table management, access controls, and data lineage.

Hands-on experience with Delta Live Tables (DLT) — pipeline development, expectations, and monitoring.

Strong command of Delta Lake internals — transaction log, ACID guarantees, file layout, and optimization techniques.

Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard creation.

Knowledge of Databricks Photon engine, serverless compute, and cost optimization strategies.

4+ years of PySpark development — Dataframe, Datasets, Spark SQL, RDD operations.

Expert-level SQL — window functions, lateral joins, CTEs, recursive queries, and analytical functions.

Experience with Spark performance tuning — AQE, query plans (EXPLAIN), partitioning, and caching.

Proficiency with Python for pipeline development, utilities, and automation.

Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery, Dataproc). Experience with cloud networking for Databricks: VNet/VPC injection, private endpoints, and firewall configurations.

Familiarity with IAM roles, managed identities, and service principal authentication for Databricks.

(Nice to Have):

Working knowledge of MLflow — experiment tracking, model registry, and deployment.

Experience supporting ML pipelines on Databricks for training, evaluation, and serving.

Education: UG - Any

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

NXP Semiconductors • Bengaluru Urban

On-site
INR 1,500,000 - 2,000,000
Databricks Platform Expert
Databricks Platform Expert

Michelin • Pune District

On-site
INR 1,600,000 - 2,700,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

ICICI Lombard • Mumbai

On-site
INR 2,800,000 - 4,000,000
Databricks Engineer/Lead
Databricks Engineer/Lead

Impetus Technologies • New Delhi, Pune District, Bengaluru

On-site
INR 1,800,000 - 3,000,000
Databricks Platform Expert
Databricks Platform Expert

MICHELIN France • Pune District

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

Swift Staffing • Pune District

On-site
INR 1,200,000 - 2,400,000
Databricks Platform Expert
Databricks Platform Expert

Michelin • Maharashtra

On-site
INR 2,400,000 - 4,800,000
Databricks - Data Engineer
Databricks - Data Engineer

Tredence Inc. • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Databricks Platform Expert
Databricks Platform Expert

Michelin Reifenwerke AG & Co. KGaA • Pune District

On-site
INR 1,200,000 - 1,800,000
26-2292: Databricks Engineer- India, Hyderabad
26-2292: Databricks Engineer- India, Hyderabad

Navitas Business Consulting, Inc. • NTR

On-site
INR 2,500,000 - 4,500,000