Senior Linguist (Text LLM – Model Evaluation)

BharatGen

Mumbai

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A leading tech firm in Mumbai is seeking a Text LLM Model Evaluation Lead to oversee the evaluation processes of their language models. You will design evaluation frameworks and collaborate with engineers and linguists to ensure quality assessments across tasks and languages. Ideal candidates have a Master's or PhD with strong experience in NLP, proficiency in Python, and a passion for refining evaluation methodologies. This role contributes significantly to the advancement of multilingual NLP capabilities.

Qualifications

  • 3+ years of experience in NLP or GenAI projects with evaluation exposure.
  • Experience managing multi-language annotation or evaluation projects is preferred.
  • Familiarity with bias detection, toxicity analysis, or fairness evaluation.

Responsibilities

  • Design and manage evaluation frameworks for text-based large language models.
  • Conduct error and trend analysis on model outputs.
  • Train and mentor junior linguists in evaluation techniques.

Skills

Human evaluation framework design
Analytical reporting
Collaboration with ML teams
Multilingual evaluation projects
Bias detection awareness

Education

Master's or PhD in Linguistics/Computational Linguistics

Tools

Python
LangGraph
DSPy
AutoGen
CrewAI

Job description

Job Summary

The Text LLM Model Evaluation Lead will own the end-to-end process of evaluating BharatGen’s text-based large language models. You will design human evaluation frameworks, test sets, rubrics, and metrics that assess model outputs across multiple tasks and languages. Working closely with ML engineers, linguists, and data operations, you’ll ensure that every model iteration is measured with rigor, fairness, and linguistic precision.

Key Responsibilities
  • Design and manage evaluation frameworks for BharatGen’s Text LLM, covering diverse tasks such as summarization, dialogue, question answering, reasoning.
  • Define evaluation dimensions (coherence, factuality etc.).
  • Develop human evaluation rubrics, task-specific test sets for multiple languages.
  • Establish evaluation workflows using human judgment that complement automated metrics.
  • Create documentation, checklists, and SOPs to ensure replicability of evaluations across model versions.
  • Collaborate with the Data Ops Manager to execute large-scale human evaluations across languages, align throughput, timelines, and cost controls.
  • Review and refine annotation guidelines to ensure inter-annotator consistency.
  • Design sampling and spot-checking methods for maintaining high data integrity.
  • Implement inter-annotator agreement tracking, quality audits.
  • Analytical Evaluation & Reporting:
  • Conduct error and trend analysis on model outputs across evaluation rounds.
  • Interpret results to highlight strengths, regressions, or recurring weaknesses.
  • Present findings and recommendations to ML engineers and leadership in structured, data-backed reports.
  • Collaborate with the ML team to refine models based on evaluation results and feedback loops.
  • Metrics & Tooling:
  • Identify or adapt suitable automatic evaluation metrics (e.g., BLEU, ROUGE, BERTScore, toxicity classifiers, etc.) to complement human evaluation.
  • Use simple scripts/dashboards to track scores, trends, & evaluation throughput.
  • Partner with Data Ops and product engineers to improve internal tools for managing evaluation tasks and results visualization.
  • Cross-functional Collaboration:
  • Train and mentor junior linguists in designing high-quality evaluation schemes.
  • Participate in design reviews to ensure evaluation insights are integrated into model training and product goals.
  • Master’s or PhD in Linguistics/Computational Linguistics with 3+ years of experience working on NLP or GenAI projects, with exposure to model evaluation, test-set design, or linguistic quality assessment.
  • Proficiency in Python and agentic frameworks such as LangGraph, DSPy, AutoGen, CrewAI, etc.
  • Experience collaborating with ML or data science teams on evaluation or model analysis workflows.
  • Experience managing multi-language annotation or evaluation projects is preferred.
  • Experience with multilingual LLMs or Indian language NLP.
  • Exposure to instruction-tuning, safety evaluation, or RLHF workflows.
  • Familiarity with bias detection, toxicity analysis, or fairness evaluation.
  • Prior experience training or mentoring annotation/evaluation teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Linguist (Speech)
Senior Linguist (Speech)

BharatGen • Mumbai

On-site
INR 1,000,000 - 1,400,000
Linguistic Data Operations Manager
Linguistic Data Operations Manager

BharatGen • Mumbai

On-site
INR 1,500,000 - 2,500,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

The LHR Group • Dadri

On-site
INR 4,000,000 - 8,000,000
AI Model Governance Specialist (Risk)
AI Model Governance Specialist (Risk)

ICICI Securities • Navi Mumbai

On-site
INR 3,000,000 - 6,000,000
LLM Model Developer
LLM Model Developer

Accenture in India • Maharashtra

On-site
INR 1,200,000 - 2,400,000
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)

Mindtel • Dadri

On-site
INR 900,000 - 1,300,000
On-site innovation lab in India
Flexible hours
ML Researcher — NLP
ML Researcher — NLP

RippleWorks Inc. • India

On-site
INR 1,500,000 - 2,200,000
Language Resource Manager
Language Resource Manager

TIH | IIT Bombay • Mumbai Suburban

On-site
INR 800,000 - 1,200,000
Large Language Model Architect – 18 years
Large Language Model Architect – 18 years

HypTechie • Pune District

On-site
INR 4,000,000 - 7,000,000
Competitive compensation and benefits
AI Developer - Large Language Models
AI Developer - Large Language Models

Volody • Maharashtra

On-site
INR 1,200,000 - 1,800,000