Data Scientist

The Nielsen Company

Bengaluru

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Nielsen Company in Bengaluru is seeking a Hybrid Data Scientist to design and manage end-to-end data pipelines for Incremental Reach and Audience Measurement.

You will implement Bayesian and ML models, deploy them in production, and collaborate with cross-functional teams to quantify the lift of Digital media over Linear TV baselines.

Required 3–6 years of experience in statistical modeling with Python and SQL, plus exposure to PySpark or Dask; knowledge of cloud platforms (AWS/GCP) is a plus.

Qualifications

  • 3–6 years in statistical model development with Python and SQL.
  • Experience with distributed computing (PySpark or Dask) is a plus.
  • Strong knowledge of TV/media analytics concepts.
  • Bachelor's or Master's in quantitative field or equivalent experience.

Responsibilities

  • Develop end-to-end data pipelines in Python to process Linear TV and Digital ad data.
  • Implement Bayesian models and GBMs to quantify lift and drivers.
  • Productionize models via APIs or containers and manage cloud data warehouses.
  • Collaborate with stakeholders to calibrate cross-media measurements.

Skills

Python
SQL
PySpark
Dask
XGBoost
LightGBM
Bayesian Methods
Statistics
Data Pipelines
Model Deployment

Education

Bachelor's or Master's in a quantitative field

Tools

Docker
Airflow
Snowflake
AWS
GCP

Job description

Job Description
Role Overview

As a Hybrid Data Scientist you will sit at the intersection of high-scale data pipelining and advanced statistical methodology. You will be responsible for the end-to-end lifecycle of Incremental Reach and Audience Measurement products—from architecting Python-based data pipelines to implementing sophisticated Bayesian and Machine Learning models that quantify the lift of Digital media over a Linear TV baseline.

Key Responsibilities
  • Advanced Statistical Modeling (The 'Science' Side)
    • Incremental Reach Frameworks: small-N datasets: implement Bayesian Model Averaging (BMA) to cycle through regression combinations, providing robust coefficients and credible intervals when study data is limited.
    • Large-Scale Prediction: deploy Gradient Boosted Regression Trees (GBM) to identify non-linear patterns and rank the impact of 'Reach Drivers' (Media Weight, On-Target %, Frequency).
    • Audience Deduplication: use Maximum Entropy (MaxEnt) models to estimate unique audience reach across fragmented platforms by reconciling census and panel data.
    • Mixed-Effect Models: use Hierarchical/Multilevel modeling to account for nested data (e.g., campaigns nested within specific industry verticals).
    • Causal Lift: apply Synthetic Control Methods to measure incremental shifts in behavior for campaigns with fixed timeframes where a clean control group is unavailable.
  • Data Engineering & Pipeline Architecture (The 'Engineering' Side)
    • Python-Centric ETL: architect and maintain robust data pipelines using Python (Pandas, PySpark) to ingest, clean, and harmonize data from Linear TV logs and Digital ad servers.
    • Feature Engineering: automate the extraction of Base Drivers (GRP, Reach Efficiency, Seasonality) and Custom Drivers (Share of Voice, Flighting) into a supervised learning-ready schema.
    • Productionization: wrap statistical models into production-grade APIs or scheduled containers (Docker/Airflow) to ensure repeatable and scalable measurement.
    • Cloud Operations: manage large-scale datasets within Cloud Data Warehouses (Snowflake, AWS, or GCP), optimizing SQL queries for high-performance analytics.
    • Control/Test Logistics: design scientifically valid Control and Test groups, ensuring proper randomization or using Propensity Score Matching to mitigate selection bias.
    • Variable Importance: provide stakeholders with Posterior Inclusion Probabilities to identify which media levers (Duration, Weight, etc.) most consistently drive incremental reach.
    • Cross-Media Calibration: reconcile Linear TV's 'One-to-Many' metrics with Digital's 'One-to-One' tracking to provide a unified view of the consumer.
Qualifications
  • Experience: 3-6 years of statistical model development and mastery of Python (specifically for data manipulation and ML) and advanced SQL. Experience with PySpark or Dask for distributed computing is a plus.
  • Statistical Mastery: proven experience with GBM (XGBoost/LightGBM) and Bayesian Frameworks (e.g., PyMC, Stan, or R-BMA) among other Data Science models.
  • Media Knowledge: understanding of Linear TV vs. Digital dynamics, including Reach/Frequency, GRPs, and Deduplication logic.
  • Education: Bachelor's or Master's in a quantitative field (Statistics, Computer Science, Economics) or equivalent professional experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist- Nielsen
Data Scientist- Nielsen

The iScale • Bengaluru

Hybrid
INR 1,800,000 - 2,400,000
Senior Data Scientist – Part-Time
Senior Data Scientist – Part-Time

Jobtailor • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Data Scientist
Data Scientist

Nielsen • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Senior Data Scientist
Senior Data Scientist

Jobtailor • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Data Scientist- CSA
Data Scientist- CSA

PivotRoots • Mumbai

On-site
INR 1,500,000 - 2,600,000
AI / ML Data Scientist I
AI / ML Data Scientist I

Thenielsencompany • Bengaluru

On-site
INR 600,000 - 1,100,000
Senior Data Scientist II
Senior Data Scientist II

Nielsen • Bengaluru

On-site
INR 3,000,000 - 7,500,000
Senior Data Scientist
Senior Data Scientist

Durapid Technologies Pvt Ltd • India

On-site
INR 4,000,000 - 7,000,000
Assistant Manager-Data Science-Data Scientist
Assistant Manager-Data Science-Data Scientist

EXL • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Data Scientist
Data Scientist

Ultra Platform • Gurugram District

On-site
INR 2,000,000 - 4,500,000