Research Fellowship: Agent Intelligence & Evaluation

Ixigo

Gurugram District

Vor Ort

INR 2.500.000 - 6.000.000

Vollzeit

Vor 5 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Ixigo is seeking a research-focused professional to advance voice and dialogue systems. You will design evaluation frameworks, craft adversarial datasets across accents, and contribute to end-to-end observability of conversational interactions.

Ideal candidates hold a PhD in ML/NLP or related fields, are proficient in Python, and know speech models such as Whisper or Conformer, plus observability tools like OpenTelemetry or Langfuse.

Qualifikationen

  • PhD candidates or equivalent track record in ML/NLP/speech research.
  • Strong programming in Python and research-driven mindset.
  • Familiarity with speech models, LLM tool-use, or observability stacks.

Aufgaben

  • Design evaluation frameworks for voice and dialogue systems.
  • Build datasets across accents and edge cases for robust testing.
  • Develop end-to-end observability across audio, STT, LLM, and TTS paths.

Kenntnisse

Python
ML/NLP intuition
Speech systems familiarity
Research mindset

Ausbildung

PhD in ML/NLP/Speech

Tools

Whisper
Conformer
OpenTelemetry
Langfuse
Arize
Hamming

Jobbeschreibung

Job Description

Voice agents fail in ways traditional software doesnt. An ASR confidence drop on a regional accent misfires a tool call, an LLM hallucinates a policy because upstream latency broke turn-taking, and support teams roll these agents back within a week without anyone able to explain what went wrong.

What youll work on

Over 4 months, youll take on one or two of the following, shaped by your interests.

Evaluation frameworks.

Text-only evals miss most of what matters in voice: barge-in, prosody, latency-induced errors, cross-turn context loss. Youll design audio-native metrics, generate adversarial conversational datasets across accents and edge cases, and build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures.

End-to-end observability.

Tracing a failed interaction means correlating audio packets, STT hypotheses, LLM reasoning traces, tool calls, and TTS output back to a single conversation ID. Youll help shape the schema and analysis layer that makes cascade failures visible across the stack.

Self-improvement systems.

Once you can measure and trace, the interesting work is closing the loop: mining production traces for failure patterns, generating targeted fine-tuning data or prompt updates, and validating that fixes hold under adversarial replay.

Qualifications
Who were looking for

Someone who cares about the research questions for their own sake, and equally cares whether the work ships. Papers at Interspeech, ACL, NeurIPS, or EMNLP on speech, dialogue systems, agent evaluation, or human-AI interaction are directly relevant.

Comfortable in Python, and familiar with at least one of: speech models (Whisper, Conformer variants), LLM tool-use and agent frameworks, or observability stacks (OpenTelemetry, Langfuse, Arize, Hamming). Current PhD students in ML, NLP, or speech are the strong default; exceptional MS students or research engineers with a publication track record are welcome to apply.

Nice to have

Prior work on evaluation methodology, dataset synthesis, or interpretability. Experience with real-time systems, telephony, or streaming pipelines. A blog, repo, or workshop paper that shows how you think in public.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Research Fellowship: Agent Intelligence & Evaluation
Research Fellowship: Agent Intelligence & Evaluation

SmartRecruiters, Inc. • New Delhi

Vor Ort
INR 508.000 - 608.000
Mentorship"
Co-authorship
Access to enterprise data
Research Fellowship: Agent Intelligence & Evaluation
Research Fellowship: Agent Intelligence & Evaluation

Ixigo • Delhi

Vor Ort
INR 335.000 - 558.000
Research Engineer - Agent Intelligence & Evaluation
Research Engineer - Agent Intelligence & Evaluation

ixigo • Delhi

Vor Ort
INR 1.400.000 - 2.000.000
Meaningful equity
Autonomy over tooling
Support to publish
Research Engineer - Agent Intelligence & Evaluation
Research Engineer - Agent Intelligence & Evaluation

SmartRecruiters, Inc. • New Delhi

Vor Ort
INR 1.800.000 - 2.800.000
Equity
Applied AI Engineer
Applied AI Engineer

Synth (YC S21) • Bengaluru

Vor Ort
INR 1.500.000 - 2.500.000
Senior Engineer — ASR / TTS / Speech LLM (Training + Eval + Integration)
Senior Engineer — ASR / TTS / Speech LLM (Training + Eval + Integration)

OutcomesAI • Bengaluru

Vor Ort
INR 2.500.000 - 4.200.000
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)
AI/LLM Speech & Audio AI Evaluation Specialist (Int. Voice)

Mindtel • Dadri

Vor Ort
INR 900.000 - 1.300.000
On-site innovation lab in India
Flexible hours
Senior Research Scientist, Agent Evaluation
Senior Research Scientist, Agent Evaluation

Snow Planet • Hyderabad, Ahmedabad District

Vor Ort
INR 2.400.000 - 4.200.000
AI Agents Applied Research/Engineering Lead - Executive Director
AI Agents Applied Research/Engineering Lead - Executive Director

JPMorgan Chase & Co. • Bengaluru

Hybrid
INR 6.000.000 - 12.000.000
Principal / Staff Software Engineer
Principal / Staff Software Engineer

Nykaa • Bengaluru

Vor Ort
INR 2.500.000 - 6.000.000