Senior Databricks Engineer

Ex

Pune District

On-site

INR 4,000,000 - 7,000,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

EXL is seeking an experienced Databricks Platform Engineer to design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod. You will configure Unity Catalog for governance and optimize Delta Lake patterns for robust lakehouse architectures.

You will develop scalable PySpark/SQL data pipelines, implement Delta Live Tables, and support MLflow workflows for AI/ML workloads in a cloud‑integrated environment.

Qualifications

  • 9–12 years of experience in data engineering / platform engineering.
  • Hands‑on with Databricks platform including Unity Catalog and Delta Lake.
  • Experience building and optimizing PySpark/SQL data pipelines and Delta Lake architectures.

Responsibilities

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod.
  • Configure Unity Catalog for data governance and access control.
  • Develop Delta Lake tables with partitioning, Z‑ordering, and file compaction; implement Delta Live Tables pipelines.

Skills

Databricks
PySpark
Delta Lake
SQL
Unity Catalog
Cloud DevOps
MLflow
Airflow

Tools

Azure DevOps
GitHub Actions
Terraform
Auto Loader
Kafka

Job description

  • Experience (In Years) 9-12
Job Description

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine‑grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto‑scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace‑level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi‑task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z‑ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built‑in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real‑time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead‑letter queue patterns in production pipelines.

MLflow & AI/ML Workloads

  • Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
  • Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
  • Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
  • Enable GenAI workloads — LLM fine‑tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
  • Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.

Cloud Integration & DevOps

  • Integrate Databricks with cloud‑native services: Azure Data Lake Storage (ADLS).
  • Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
  • Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure‑as‑code (IaC) deployment of Databricks resources.
  • Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
  • Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.

Governance, Security & Optimization

  • Implement row‑level security, column masking, and dynamic data views using Unity Catalog policies.
  • Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
  • Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
  • Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
  • Document architecture decisions, runbooks, and operational guides for Databricks workloads.
Responsibilities

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine‑grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto‑scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace‑level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi‑task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z‑ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built‑in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real‑time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead‑letter queue patterns in production pipelines.

MLflow & AI/ML Workloads

  • Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
  • Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
  • Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
  • Enable GenAI workloads — LLM fine‑tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
  • Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.

Cloud Integration & DevOps

  • Integrate Databricks with cloud‑native services: Azure Data Lake Storage (ADLS).
  • Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
  • Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure‑as‑code (IaC) deployment of Databricks resources.
  • Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
  • Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.

Governance, Security & Optimization

  • Implement row‑level security, column masking, and dynamic data views using Unity Catalog policies.
  • Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
  • Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
  • Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
  • Document architecture decisions, runbooks, and operational guides for Databricks workloads.
Qualifications

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine‑grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto‑scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace‑level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi‑task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z‑ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built‑in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real‑time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead‑letter queue patterns in production pipelines.

MLflow & AI/ML Workloads

  • Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
  • Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
  • Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
  • Enable GenAI workloads — LLM fine‑tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
  • Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.

Cloud Integration & DevOps

  • Integrate Databricks with cloud‑native services: Azure Data Lake Storage (ADLS).
  • Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
  • Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure‑as‑code (IaC) deployment of Databricks resources.
  • Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
  • Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.

Governance, Security & Optimization

  • Implement row‑level security, column masking, and dynamic data views using Unity Catalog policies.
  • Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
  • Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
  • Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
  • Document architecture decisions, runbooks, and operational guides for Databricks workloads.
About Us

EXL (NASDAQ: EXLS) is a leading data analytics and digital operations and solutions company. We partner with clients using a data and AI‑led approach to reinvent business models, drive better business outcomes and unlock growth with speed. EXL harnesses the power of data, analytics, AI, and deep industry knowledge to transform operations for the world’s leading corporations in industries including insurance, healthcare, banking and financial services, media and retail, among others. EXL was founded in 1999 with the core values of innovation, collaboration, excellence, integrity and respect. We are headquartered in New York and have more than 54,000 employees spanning six continents. For more information, visit www.exlservice.com .

EXL never requires or asks for fees/payments or credit card or bank details during any phase of the recruitment or hiring process and has not authorized any agencies or partners to collect any fee or payment from prospective candidates. EXL will only extend a job offer after a candidate has gone through a formal interview process with members of EXL’s Human Resources team, as well as our hiring managers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Databricks Architect
Databricks Architect

EXL • Pune District

On-site
INR 1,500,000 - 2,800,000
Sr Databricks Engineer
Sr Databricks Engineer

EXL • Pune District

Hybrid
INR 3,000,000 - 5,500,000
Assistant Manager
Assistant Manager

EXL • Dadri

On-site
INR 350,000 - 550,000
3726774-Lead Assistant Manager
3726774-Lead Assistant Manager

EXL • Gurugram District

On-site
INR 1,500,000 - 2,200,000
Mentoring program
Analytics training
Guidance and coaching
1984760-Assistant Vice President
1984760-Assistant Vice President

EXL • Gurugram District

On-site
INR 3,500,000 - 6,500,000
AVP - EPM
AVP - EPM

Ex • Dadri

On-site
INR 1,800,000 - 2,400,000
Senior Manager / AVP - Analytics and Insights
Senior Manager / AVP - Analytics and Insights

EXL • Gurugram District

On-site
INR 2,500,000 - 4,500,000
Mentoring program
Professional development
Competitive compensation
1984866-Manager
1984866-Manager

EXL • Gurugram District

On-site
INR 1,500,000 - 2,100,000
Executives
Executives

Ex • Pune District

On-site
INR 300,000 - 500,000
2067748-Lead Assistant Manager
2067748-Lead Assistant Manager

EXL • Gurugram District

On-site
INR 900,000 - 1,500,000