Embedded Data Scientist, Chanakya

Sarvam AI

Delhi

On-site

INR 800,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sarvam AI in Delhi is looking for an Embedded Data Scientist to transform complex client data into semantic structures for AI systems. The role involves working with diverse datasets, designing ontologies, and developing data ingestion pipelines to improve AI performance.

The ideal candidate will have 2–5 years of experience in data science with strong Python skills, solid grounding in machine learning, and familiarity with LLM-based systems. A dynamic work environment awaits where you'll impact real-world applications of AI.

Qualifications

  • 2–5 years in data science, applied machine learning, or large-scale data analysis roles.
  • Solid grounding in machine learning fundamentals.
  • Experience designing or working with data schemas or semantic data structures.

Responsibilities

  • Understand client's data landscape.
  • Design domain ontologies for data structuring.
  • Collaborate with teams to translate semantic structures.

Skills

Data science expertise
Strong Python skills
ML fundamentals
Experience with unstructured datasets
Familiarity with LLM-based systems

Tools

pandas
NumPy

Job description

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full‑stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

Embedded Data Scientists transform complex client data into structures that AI systems can reliably reason over. You are deployed alongside Strategic Deployment Engineers at client sites, working directly with client data environments to understand, structure, and operationalise large‑scale datasets.

This means working with heterogeneous, multimodal data – including documents, images, audio, geospatial data, and structured records – and designing the semantic structures that allow AI systems to interpret and reason over that data.

What You'll Do
  • Understand the client's data landscape across documents, imagery, audio, geospatial data, and structured records – including data sources, formats, workflows, and domain terminology
  • Design domain ontologies representing entities, relationships, hierarchies, and operational concepts within the client's data environment
  • Define document segmentation and chunking strategies that preserve semantic meaning and support effective retrieval
  • Work with heterogeneous datasets and define how different modalities should be indexed, embedded, and linked
  • Collaborate with Strategic Deployment Engineers to translate semantic structures into operational data ingestion pipelines
  • Evaluate how well the AI system retrieves and reasons over client data, and refine structures to improve performance
  • Collaborate with the models and other teams to define benchmarks and evaluation criteria that reflect real‑world deployment conditions
  • Translate insights from client data environments into structured signals for product and engineering teams
What We're Looking For
  • 2–5 years in data science, applied machine learning, or large‑scale data analysis roles
  • Strong Python skills including pandas, NumPy, and modern NLP or LLM tooling
  • Solid grounding in ML fundamentals – enough to understand model behaviour, contribute to evaluation design, and collaborate with a models team on training and benchmarking
  • Experience working with large unstructured datasets including documents, transcripts, reports, or operational records
  • Familiarity with LLM‑based systems, retrieval pipelines, or vector search systems
  • Experience designing or working with data schemas, metadata frameworks, entity models, or semantic data structures
Signals We Look For
  • You've worked with real‑world, messy, unstructured data and built something rigorous from it
  • You are comfortable designing structure where none exists – defining schemas, ontologies, and metadata frameworks from scratch
  • You can translate complex data insights into explanations that engineers and client stakeholders can act on
Who You Are

You are comfortable operating with autonomy in client environments; you don't need a data team around you to do rigorous work. You move fluently between domain understanding, data modelling, and AI system design. You move between the technical and the operational: you understand what the data means in the context of what operators actually do with it.

Bonus Points
  • Familiarity with knowledge graphs, ontologies, or semantic data modelling
  • Experience with multimodal datasets (text, imagery, audio, geospatial, or structured data)
  • Experience operating in constrained or air‑gapped environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Sarvam • Delhi

On-site
INR 1,000,000 - 2,000,000
Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Neara • Delhi

On-site
INR 1,000,000 - 1,500,000
Embedded Infrastructure Engineer, Chanakya
Embedded Infrastructure Engineer, Chanakya

Neara • Delhi

On-site
INR 1,200,000 - 2,000,000
Impactful work
High ownership
Collaborative team environment
Embedded Infrastructure Engineer, Chanakya
Embedded Infrastructure Engineer, Chanakya

Sarvam • Delhi

On-site
INR 1,400,000 - 2,000,000
Strategic Deployment Engineer, Chanakya
Strategic Deployment Engineer, Chanakya

Neara • Delhi

On-site
INR 1,200,000 - 1,800,000
ML Engineer (Data), Foundational Models
ML Engineer (Data), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,000,000
High ownership and impact
Work alongside top talent
AI-first environment
Forward Deployed Engineer
Forward Deployed Engineer

Sarvam AI • Bengaluru

On-site
INR 1,800,000 - 3,000,000
ML Ops Engineer, Chanakya
ML Ops Engineer, Chanakya

Sarvam • Delhi

On-site
INR 1,200,000 - 2,000,000
High ownership
High impact from day one
AI-first approach
Data Scientist - Evaluations, Chanakya
Data Scientist - Evaluations, Chanakya

Neara • India

On-site
INR 1,200,000 - 1,800,000
High ownership and impact
AI-first working environment
Collaborative team of experts
ML Ops Engineer, Chanakya
ML Ops Engineer, Chanakya

Neara • Delhi

On-site
INR 1,500,000 - 2,000,000
High ownership in projects
Collaborative team environment
Opportunity to work on impactful AI solutions