Data Scientist Staff

Stryker

Gurugram District

On-site

INR 3,600,000 - 6,000,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stryker is seeking an experienced data engineer to build and own end-to-end data pipelines on Azure Databricks, from ingestion to curated Gold datasets. You will manage CI/CD in Azure DevOps and design the Power BI consumption layer over lakehouse data.

Responsibilities include data quality, lineage, regulatory validation (FDA 21 CFR Part 11, EU MDR), and developing ML/GenAI solutions to deliver analytics-ready insights. 7–10 years of experience required; on-site in Gurugram region.

Qualifications

  • 7–10 years of data engineering experience with Azure Databricks.
  • Proficient in Delta Lake, PySpark and Unity Catalog.
  • Strong SQL and ETL/ELT design skills for large datasets.

Responsibilities

  • Build end-to-end data pipelines on Azure Databricks with medallion architecture.
  • Own the delivery pipeline: CI/CD in Azure DevOps across envs.
  • Design the Power BI semantic models and deployment pipelines.
  • Implement data quality, validation and schema-drift controls.
  • Maintain lineage, audits per FDA 21 CFR Part 11 and EU MDR.
  • Develop ML/GenAI solutions for analytics-ready insights.
  • Monitor and cost-optimize production workloads and tuning.
  • Collaborate with platform team; translate requirements and assess impact.

Skills

PySpark
Delta Lake
Azure Databricks
Spark tuning
Unity Catalog
SQL
ETL/ELT design
Python
GenAI/ML
Data governance

Tools

Power BI
Azure DevOps
Git
CI/CD pipelines
Azure DL/Storage

Job description

What you will do:
  • Build and own end-to-end data pipelines on Azure Databricks: from ingestion through Bronze/Silver/Gold medallion transformation to curated Gold datasets, including incremental and historical loads from varied sources.
  • Own the delivery pipeline: repository structure, branching strategy, and YAML-based CI/CD in Azure DevOps and manage promotion of code, data and reports across Dev, QA, UAT and Production.
  • Design and maintain the Power BI consumption layer: semantic models, datasets and reports over lakehouse data including workspace promotion and capacity management.
  • Design and implement data quality, validation and reconciliation frameworks: proving that automated output matches the manual baseline it replaces and handling schema drift without silent failure.
  • Maintain lineage, audit trails and validation evidence: to the standard required by FDA 21 CFR Part 11 and EU MDR, and contribute to test strategy and UAT execution alongside business SMEs.
  • Develop and deploy machine learning and Generative AI solutions: to develop analytics-ready insight for various use cases and business objectives.
  • Monitor, tune and cost-optimise production workloads: cluster sizing and autoscaling, Spark and Delta performance tuning, alerting, and incident triage when pipelines fail.
  • Partner with the central platform team to specify infrastructure requirements precisely, diagnose provisioning gaps, and unblock dependencies before they reach the critical path. Translate business and regulatory requirements into scoped solutions with SMEs and stakeholders, quantify the financial impact, and present findings to technical and non-technical audiences. Contribute to the platform's evolution toward Microsoft Fabric / OneLake, assessing what transfers cleanly and what must be rebuilt.
What you need:

Must have skills

  • 7-10 years of experience. Hands-on Azure Databricks depth: Delta Lake, medallion architecture, PySpark, cluster and job management, Unity Catalog, and performance optimisation (partitioning, caching, broadcast joins, AQE, OPTIMIZE / Z-ORDER).
  • The wider Azure data stack: ADLS Gen2, Azure Data Factory, Synapse, with strong SQL and ETL/ELT design including incremental ingestion patterns. SQL Server and Key Vault exposure expected.
  • Unstructured and siloed data sourcing: ingestion and remediation of unstructured content from distributed sources (scanned or flattened documents, Excel-to-PDF snapshots that resist standard OCR, SharePoint repositories, attachments) using OCR, document intelligence and AI techniques. CI/CD and environment ownership: Git, Azure DevOps, YAML pipelines, secrets and service principal management, and multi-environment promotion across Dev/QA/UAT/Prod.
  • Power BI to a build-and-own standard: semantic models, Power Query/M, DAX, and deployment pipelines, including where a direct lake-to-BI connection is unavailable.
  • Expert Python and applied machine learning: at least two of time series forecasting, classification, computer vision or predictive modelling, plus practical LLM/GenAI experience.
  • Production data quality and governance discipline: validation frameworks, data lineage, audit logging, and role-based security. Demonstrable hands-on delivery ownership: within a centrally governed enterprise platform, with the financial acumen to size and defend the business impact of what you build.
Preferred skills
  • Regulated-environment exposure: medical devices, pharmaceuticals, healthcare or manufacturing, with familiarity with validation, audit or compliance lifecycles (FDA 21 CFR Part 11, EU MDR, GxP).
  • Microsoft Fabric / OneLake experience: assessed as transition capability, not a gate. Applied NLP beyond document processing: embeddings, semantic similarity, fuzzy matching and clustering for text standardisation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Baker Tilly • Bengaluru

On-site
INR 1,200,000 - 3,000,000
Data Engineer
Data Engineer

Kumaran Systems • Chennai District

On-site
INR 4,000,000 - 7,000,000
Data Engineering Lead
Data Engineering Lead

Kumaran Systems • Hyderabad

On-site
INR 2,800,000 - 4,000,000
Data Engineer
Data Engineer

Kumaran Systems • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Finarb • Kolkata District

On-site
INR 1,200,000 - 1,800,000
Sr. Azure Data Engineer
Sr. Azure Data Engineer

Noventiq • Dadri

On-site
INR 1,200,000 - 1,800,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineering Manager with Azure & Databricks
Data Engineering Manager with Azure & Databricks

PwC • Gurugram District, Bengaluru, Hyderabad

Hybrid
INR 4,200,000 - 6,800,000
Databricks Engineer
Databricks Engineer

EXL • Pune District

On-site
INR 3,500,000 - 5,500,000