Data Scientist - Evaluations, Chanakya

Neara

India

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High ownership and impact
AI-first working environment
Collaborative team of experts

Job summary

A technology company is seeking a Data Scientist for Evaluations in Bengaluru. The role involves designing evaluation frameworks for AI outputs, defining quality metrics, and collaborating with domain experts. Candidates should have 3–6 years of experience in data science, strong statistical knowledge, and Python proficiency. Experience with LLMs and evaluating unstructured domain data is essential. This position offers high ownership and impact in AI development.

Qualifications

  • 3–6 years in data science or applied AI, including 2 years with LLMs.
  • Strong statistics and probability fundamentals.
  • Experience designing evaluation frameworks from scratch.

Responsibilities

  • Design and build evaluation frameworks for AI outputs.
  • Define quality metrics with domain experts.
  • Run structured evaluation cycles pre- and post-deployment.

Skills

Data science
Machine learning
Statistical analysis
Python programming
Prompt engineering

Tools

pandas
NumPy
HuggingFace datasets
LangSmith

Job description

Data Scientist - Evaluations

Job type: Full Time · Department: Engineering · Work type: On‑Site · Location: Bengaluru, Karnataka, India

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full‑stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

This is an intellectually demanding role. The Data Scientist anchors the evaluations function for this vertical — designing, building, and maintaining evaluation frameworks that measure model and system quality in operational context. You are not running standard benchmarks. You are building domain‑specific eval harnesses for high‑stakes use cases where a wrong answer carries real consequences.

You will work closely with the MLOps Engineer, the PM, and the deployment team. The evaluations you design are the mechanism by which the team determines whether what we've built is good enough to deploy — and whether it stays good after deployment.

What You'll Do
  • Design and build evaluation frameworks for Sarvam's AI outputs across domain‑specific requirements: document comprehension, command summarisation, geospatial reasoning, enterprise workflow automation, and others as they emerge.
  • Define quality metrics in collaboration with domain experts and clients; translate operational requirements into measurable, defensible signals.
  • Run structured evaluation cycles pre‑ and post‑deployment; build dashboards that surface model quality in production.
  • Identify failure modes, edge cases, and distribution shifts — with the bias of someone looking for what's wrong, not confirming what's right.
  • Collaborate with the MLOps Engineer to operationalise eval pipelines — automated, triggered by deployment events, versioned, and reproducible.
  • Build and manage domain‑specific datasets for fine‑tuning, evaluation, and benchmarking — including human annotation workflows where needed.
  • Publish internal findings and quality reports that feed the product and engineering roadmap.
What We're Looking For
  • 3–6 years in data science, ML research, or applied AI; at least 2 years working with LLMs in production contexts.
  • Strong statistics and probability fundamentals — you understand what makes an evaluation valid and what makes it misleading.
  • Experience designing evaluation frameworks from scratch: custom metrics, inter‑rater reliability, red‑teaming methodologies.
  • Python proficiency; comfort with pandas, NumPy, HuggingFace datasets, RAGAS, EleutherAI Eval Harness, LangSmith, or equivalent.
  • Experience with prompt engineering, model fine‑tuning, or RLHF in applied settings.
  • Ability to work with unstructured domain data: PDFs, doctrine documents, transcripts, and field reports.
Bonus Points
  • Prior work in high‑stakes domains (healthcare, legal, defence, finance) where output quality carries real‑world consequences.
  • Red‑teaming or adversarial evaluation experience.

Note: We are looking for people who can own the outcomes described here, not people who match every line of this specification. If this problem excites you and you believe you can do this work, we want to hear from you.

Why Sarvam?
  • Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar.
  • High ownership and high impact, from day one.
  • Everything we do is AI‑first, from the way we build and ship to the way we think about problems.
  • You can work on problems that could change how an entire country learns, works, and communicates.

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Neara • Delhi

On-site
INR 1,000,000 - 1,500,000
Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Sarvam • Delhi

On-site
INR 1,000,000 - 2,000,000
ML Engineer (Data), Foundational Models
ML Engineer (Data), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,000,000
High ownership and impact
Work alongside top talent
AI-first environment
ML Ops Engineer, Chanakya
ML Ops Engineer, Chanakya

Neara • Delhi

On-site
INR 1,500,000 - 2,000,000
High ownership in projects
Collaborative team environment
Opportunity to work on impactful AI solutions
Strategic Deployment Engineer, Chanakya
Strategic Deployment Engineer, Chanakya

Neara • Delhi

On-site
INR 1,200,000 - 1,800,000
ML Ops Engineer, Chanakya
ML Ops Engineer, Chanakya

Sarvam • Delhi

On-site
INR 1,200,000 - 2,000,000
High ownership
High impact from day one
AI-first approach
Machine Learning Engineer, Vision
Machine Learning Engineer, Vision

Neara • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Embedded Infrastructure Engineer, Chanakya
Embedded Infrastructure Engineer, Chanakya

Neara • Delhi

On-site
INR 1,200,000 - 2,000,000
Impactful work
High ownership
Collaborative team environment
Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Sarvam AI • Delhi

On-site
INR 800,000 - 1,200,000
Backend Engineer, Chanakya
Backend Engineer, Chanakya

Neara • India

On-site
INR 1,500,000 - 2,000,000
High ownership and impact
Collaborative team environment