Research Fellowship: Agent Intelligence & Evaluation

SmartRecruiters, Inc.

New Delhi

On-site

INR 508,000 - 608,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Mentorship"
Co-authorship
Access to enterprise data

Job summary

ixigo offers a four-month research fellowship to work on self-healing voice agents for enterprise support. You will explore evaluation frameworks, end-to-end observability, and self-improvement systems, focusing on speech, LLMs, and telemetry.

The role requires Python, familiarity with Whisper/Conformer, and ML/NLP research; publications are valued. Compensation is ₹50,000 per month during the term, with mentorship and potential co-authorship.

Qualifications

  • Proficient in Python and ML/NLP research.
  • Familiar with speech models and evaluation.
  • Candidates with publications in relevant venues welcomed.

Responsibilities

  • Design audio-native evaluation frameworks for voice agents.
  • Build end-to-end observability across STT, LLM, and TTS traces.
  • Mine production traces to improve models and prompts.

Skills

Python
speech models
LLM tool-use
observability stacks

Education

PhD student in ML/NLP/speech

Tools

Whisper
Conformer
OpenTelemetry
Langfuse
Arize
Hamming

Job description

  • Full-time
Company Description

We’re building self-healing voice agents for enterprise customer support within ixigo. The system has to know when it’s failing, why it’s failing, and how to fix itself before a human notices. This fellowship sits at the intelligence layer behind that work.

Job Description

Voice agents fail in ways traditional software doesn't. An ASR confidence drop on a regional accent misfires a tool call, an LLM hallucinates a policy because upstream latency broke turn-taking, and support teams roll these agents back within a week without anyone able to explain what went wrong.

What you'll work on

Over 4 months, you'll take on one or two of the following, shaped by your interests.

Evaluation frameworks. Text-only evals miss most of what matters in voice: barge-in, prosody, latency-induced errors, cross-turn context loss. You'll design audio-native metrics, generate adversarial conversational datasets across accents and edge cases, and build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures.

End-to-end observability. Tracing a failed interaction means correlating audio packets, STT hypotheses, LLM reasoning traces, tool calls, and TTS output back to a single conversation ID. You'll help shape the schema and analysis layer that makes cascade failures visible across the stack.

Self-improvement systems. Once you can measure and trace, the interesting work is closing the loop: mining production traces for failure patterns, generating targeted fine-tuning data or prompt updates, and validating that fixes hold under adversarial replay.

Qualifications

Who we're looking for

Someone who cares about the research questions for their own sake, and equally cares whether the work ships. Papers at Interspeech, ACL, NeurIPS, or EMNLP on speech, dialogue systems, agent evaluation, or human-AI interaction are directly relevant.

Comfortable in Python, and familiar with at least one of: speech models (Whisper, Conformer variants), LLM tool-use and agent frameworks, or observability stacks (OpenTelemetry, Langfuse, Arize, Hamming). Current PhD students in ML, NLP, or speech are the strong default; exceptional MS students or research engineers with a publication track record are welcome to apply.

Nice to have

Prior work on evaluation methodology, dataset synthesis, or interpretability. Experience with real-time systems, telephony, or streaming pipelines. A blog, repo, or workshop paper that shows how you think in public.

Additional Information

What you'll get

₹50,000/month for the 4-month term, access to real enterprise conversation data under proper governance, mentorship on the research and shipping sides, co-authorship on papers that come out of the work, and a system in production running on top of what you build. Candidates are responsible for safeguarding sensitive company data against unauthorized access, use, or disclosure, and for reporting any suspected security incidents in line with the organization's ISMS (Information Security Management System) policies and procedures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer - Agent Intelligence & Evaluation
Research Engineer - Agent Intelligence & Evaluation

SmartRecruiters, Inc. • New Delhi

On-site
INR 1,800,000 - 2,800,000
Equity
Research Engineer - Agent Intelligence & Evaluation
Research Engineer - Agent Intelligence & Evaluation

ixigo • Delhi

On-site
INR 1,400,000 - 2,000,000
Meaningful equity
Autonomy over tooling
Support to publish
Research Engineer - Voice and Language AI
Research Engineer - Voice and Language AI

GreyLabs AI • Bengaluru

On-site
INR 4,000,000 - 7,000,000
ML Research Engineer Speech
ML Research Engineer Speech

Blue Machines AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
ML Research Engineer, Speech
ML Research Engineer, Speech

Blue Machines AI • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior ML Research Scientist, Speech
Senior ML Research Scientist, Speech

Blue Machines AI • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Engineer — ASR / TTS / Speech LLM (Training + Eval + Integration)
Senior Engineer — ASR / TTS / Speech LLM (Training + Eval + Integration)

OutcomesAI • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Research Engineer — Agent Architectures (Coding & Autonomous Systems)
Research Engineer — Agent Architectures (Coding & Autonomous Systems)

Story Terrace Inc. • Mumbai

Remote
INR 1,800,000 - 3,200,000
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)

Mindtel • Dadri

On-site
INR 900,000 - 1,300,000
On-site innovation lab in India
Flexible hours
Lead AI Engineer
Lead AI Engineer

United States Digital Space LLC • Bengaluru

On-site
INR 3,000,000 - 5,500,000