Software Development Engineer - I (Data)

BookMyShow

Mumbai

On-site

INR 1,800,000 - 3,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BookMyShow is seeking a Data Enthusiast who blends data engineering with ML in production. You will design scalable Databricks pipelines, partner with Data Science/BI teams, and operationalize models from feature engineering to deployment and monitoring.

You will own data quality, governance, and observability while optimizing costs and performance across lakehouse architectures. This role requires collaboration with infra and platform teams.

Qualifications

  • 1-3 years of data engineering experience working on Databricks in production.
  • Strong proficiency in PySpark/Spark SQL and Python.
  • Solid understanding of Delta Lake, Unity Catalog, Lakeflow/DLT, and Databricks system tables (billing, compute, query history).
  • Experience building and maintaining feature pipelines or ML data infrastructure (feature stores, training/serving data parity).
  • Familiarity with MLflow (experiment tracking, model registry) and/or Databricks Model Serving.
  • Strong SQL skills and experience with warehouse performance tuning (query optimization, materialization strategies, cluster sizing).
  • Understanding of ML fundamentals — to discuss features, drift, and model lifecycle with Data Scientists.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using Databricks (Spark, Delta Lake, Lakeflow/DLT, Unity Catalog).
  • Develop and productionize feature pipelines for ML use cases, ensuring reliability, freshness, and reproducibility.
  • Collaborate with Data Scientists/ML Engineers to deploy models (batch and/or real-time), including via Databricks Model Serving or MLflow.
  • Optimize Spark jobs, SQL warehouses, and cluster configurations for cost and performance.
  • Build and maintain system-level observability for pipelines and ML jobs (usage, cost, quality, drift).
  • Implement data quality checks, testing, and monitoring across the medallion architecture (bronze/silver/gold).
  • Own Unity Catalog governance for datasets and features — access controls, lineage, and PII masking.
  • Partner with platform/infra teams on job orchestration, CI/CD for data & ML pipelines, and cost optimization.
  • Contribute to architecture decisions around lakehouse design, streaming vs. batch tradeoffs, and tool selection (build vs. buy).
  • Perform analysis on top of the data you build, answer ad-hoc business questions, validate metrics, and spot data quality issues before they reach stakeholders.
  • Evaluate and apply LLMs/agentic frameworks responsibly within the team, balancing accuracy, cost, and governance.

Skills

Databricks data engineering
PySpark / Spark SQL
Python
ML fundamentals
SQL & data warehouse tuning
Observability & monitoring
Feature pipelines / ML infra
Cost governance
ML model lifecycle concepts
Lakehouse concepts

Tools

Databricks
Spark
Delta Lake
Lakeflow/DLT
Unity Catalog
MLflow
Databricks Model Serving
Structured Streaming
Kafka
Airflow

Job description

The World of BookMyShow

Launched in 2007, BookMyShow, owned and operated by Big Tree Entertainment Pvt. Ltd. (founded in 1999), is India's leading entertainment destination with global operations and the one-stop shop for every entertainment need. The firm is present in over 650 towns and cities in India and works with partners across the industry to provide unmatched entertainment experiences to millions of customers.

Over the years, the company has evolved from a purely online ticketing platform for movies across 6,000 plus screens, to end-to-end management of live entertainment events including music concerts, live performances, theatricals, sports and more. Some of the key properties that BookMyShow has brought to its markets include U2's The Joshua Tree Tour, NBA's debut games in India, Disney's Aladdin, Cirque du Soleil BAZZAR as well as international artists such as Coldplay, Ed Sheeran, and Justin Bieber. BookMyShow is invested in providing the best user experience, whether on-ground or online.

The company has developed 'BookMyShow Stream', India's largest home-grown transactional video-on-demand (TVOD) platform.

Role Overview

We're looking for a Data Enthusiast who combines strong data engineering fundamentals with hands‑on experience applying machine learning in production. You'll build and scale data pipelines on Databricks, and partner closely with Data Science/Business Intelligence teams to operationalize models - from feature engineering through deployment and monitoring as well as generating insights that create business impact

Your Profile
  • Design, build, and maintain scalable ETL/ELT pipelines using Databricks (Spark, Delta Lake, Lakeflow/DLT, Unity Catalog)
  • Develop and productionize feature pipelines for ML use cases, ensuring reliability, freshness, and reproducibility
  • Collaborate with Data Scientists/ML Engineers to deploy models (batch and/or real‑time), including via Databricks Model Serving or ML flow
  • Optimize Spark jobs, SQL warehouses, and cluster configurations for cost and performance
  • Build and maintain system‑level observability for pipelines and ML jobs (usage, cost, quality, drift)
  • Implement data quality checks, testing, and monitoring across the medallion architecture (bronze/silver/gold)
  • Own Unity Catalog governance for datasets and features — access controls, lineage, and PII masking
  • Partner with platform/infra teams on job orchestration, CI/CD for data & ML pipelines, and cost optimization
  • Contribute to architecture decisions around lakehouse design, streaming vs. batch tradeoffs, and tool selection (build vs. buy)
  • Perform analysis on top of the data you build, answer ad‑hoc business questions, validate metrics, and spot data quality issues before they reach stakeholders
  • Evaluate and apply LLMs/agentic frameworks responsibly within the team, balancing accuracy, cost, and governance (e.g., row/column‑level access control on what an AI agent can query)
Your Checklist
  • 1-3 years of data engineering experience working on Databricks in production
  • Strong proficiency in PySpark/Spark SQL and Python
  • Solid understanding of Delta Lake, Unity Catalog, Lakeflow/DLT, and Databricks system tables (billing, compute, query history) Experience building and maintaining feature pipelines or ML data infrastructure (feature stores, training/serving data parity)
  • Familiarity with MLflow (experiment tracking, model registry) and/or Databricks Model Serving
  • Strong SQL skills and experience with warehouse performance tuning (query optimization, materialization strategies, cluster sizing)
  • Understanding of ML fundamentals — enough to have real conversations with Data Scientists about features, drift, and model lifecycle (you don't need to be building models yourself, but you should understand what “good” looks like)
Preferred Skills
  • Experience with real‑time/streaming architectures (Structured Streaming, Kafka, Lakebase or similar)
  • Exposure to LLM/agentic tooling (Databricks Genie, RAG pipelines, vector search)
  • Experience with cost governance/FinOps for Databricks workloads
  • Background with experimentation platforms or A/B testing infrastructure
  • Familiarity with orchestration tools (Databricks Jobs, Airflow) and CI/CD for data pipelines
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Development Engineer – I (Data)
Software Development Engineer – I (Data)

BookMyShow • Mumbai

On-site
INR 1,200,000 - 2,400,000
Databrick Data Engineer
Databrick Data Engineer

BDO India • Mumbai

On-site
INR 2,500,000 - 4,500,000
Databricks Engineer
Databricks Engineer

EXL • Pune District

On-site
INR 3,500,000 - 5,500,000
Lead Software Engineer
Lead Software Engineer

Impetus • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

ICICI Lombard • Mumbai

On-site
INR 2,800,000 - 4,000,000
Databricks Data Architect
Databricks Data Architect

Unison Group • Chennai District

On-site
INR 1,800,000 - 3,000,000
Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Databricks - Data Engineer
Databricks - Data Engineer

Tredence Inc. • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Expert Data Engineer (Databricks)
Expert Data Engineer (Databricks)

Codvo.ai • Pune District

On-site
INR 800,000 - 1,200,000
Sr. Data Engineer (Databricks)
Sr. Data Engineer (Databricks)

Blumetra Solutions • Hyderabad

Hybrid
INR 1,200,000 - 1,500,000