Research Engineer - Agent Intelligence & Evaluation

ixigo

Delhi

On-site

INR 1,400,000 - 2,000,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Meaningful equity
Autonomy over tooling
Support to publish

Job summary

ixigo is building self-healing voice agents for enterprise customer support. This fellowship owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline, and the feedback loops that let agents fix themselves.

You'll work with senior ML engineers to build robust evaluation metrics, real-time telemetry, and automated data pipelines, with equity and autonomy over tooling as part of a small, high-impact team.

Qualifications

  • 3 to 5 years in ML engineering, research engineering, or applied research.
  • Strong Python and modern ML tooling.
  • Depth in at least two of: speech and audio models, LLM agent systems, and eval or observability infrastructure.

Responsibilities

  • Evaluation infrastructure: audio-native metrics for barge-in, prosody, and turn-taking; adversarial datasets across accents and edge cases; LLM-as-judge rubrics for task success, tool-use correctness, and recovery.
  • Observability across the pipeline: tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation; analysis and alerting that surfaces cascade failures.
  • Self-improvement systems: mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions.

Skills

Python
ML tooling
Speech models
LLM agent systems
Eval infra

Tools

PyTorch
TensorFlow
Airflow

Job description

  • Full-time
Company Description

We’re building self-healing voice agents for enterprise customer support within ixigo. The system has to know when it’s failing, why it’s failing, and how to fix itself before a human notices. This fellowship sits at the intelligence layer behind that work.

Job Description

Voice agents fail in ways traditional software doesn't. ASR confidence drops on an accent and a tool call misfires. Latency breaks turn-taking and the LLM hallucinates a policy. A model swap silently regresses production and nobody catches it for a week.

We're building self-healing voice agents for enterprise customer support. This role owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline, and the feedback loops that let agents fix themselves

What you'll own
  • Evaluation infrastructure. Audio-native metrics for barge-in, prosody, and turn-taking. Adversarial datasets across accents and edge cases. LLM-as-judge rubrics for task success, tool-use correctness, and recovery.
  • Observability across the pipeline. Tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation. Analysis and alerting that surfaces cascade failures instead of hiding them.
  • Self-improvement systems. Mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions.
Qualifications

Who we're looking for 3 to 5 years in ML engineering, research engineering, or applied research. Strong Python and modern ML tooling. Depth in at least two of: speech and audio models, LLM agent systems, and eval or observability infrastructure.

Nice to have Real-time systems or telephony experience. Work on RLHF, DPO, or synthetic data pipelines. Familiarity with enterprise deployment (SOC 2, PII, data residency).

What you'll get Senior seat on a small team where research and production aren't separate orgs. Real enterprise conversation data under proper governance. Meaningful equity, autonomy over tooling, and support to publish.

Additional Information

Candidates are responsible for safeguarding sensitive company data against unauthorized access, use, or disclosure, and for reporting any suspected security incidents in line with the organization's ISMS (Information Security Management System) policies and procedures.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer — Agent Intelligence & Evaluation
Research Engineer — Agent Intelligence & Evaluation

ixigo • Gurugram District

On-site
INR 900,000 - 1,400,000
Equity
Autonomy over tooling
Governance-friendly environment
Research Engineer - Agent Intelligence & Evaluation
Research Engineer - Agent Intelligence & Evaluation

Ixigo • Gurugram District

On-site
INR 2,600,000 - 4,000,000
Research Engineer - Voice & Language AI
Research Engineer - Voice & Language AI

GreyLabs AI • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Forward Deployed AI Engineer (Voice AI Experience)
Forward Deployed AI Engineer (Voice AI Experience)

Herox • Gurugram District

On-site
INR 1,800,000 - 3,200,000
AI Engineer
AI Engineer

VoiceCare AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Research Engineer Poland +4 more
Research Engineer Poland +4 more

ElevenLabs • Bengaluru

Remote
PLN 169,000 - 255,000
Innovative culture
Growth paths
Learning & development stipend
+2
Staff AI Engineer — Agentic AI
Staff AI Engineer — Agentic AI

gnani.ai • Bengaluru

On-site
INR 4,000,000 - 12,000,000
Applied AI Engineer
Applied AI Engineer

Synth (YC S21) • Bengaluru

On-site
INR 1,500,000 - 2,500,000
AI Engineer — Agents & GenAI
AI Engineer — Agents & GenAI

Rixroent Private Limited • India

On-site
INR 1,200,000 - 2,400,000
Senior AI Engineer
Senior AI Engineer

Arcana • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000