Analyst (Speech and Audio AI evaluation)

Innodata Inc.

Uttar Pradesh

On-site

INR 600,000 - 900,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Innodata Inc. is seeking Speech & Audio AI Evaluation Specialists to assess state-of-the-art S2S, TTS, and live voice agents in our India GCC. You will judge naturalness, prosody, and cadence across global English dialects, ensuring high-quality benchmarks.

Responsibilities include blinded pairwise ratings, multi-modal validation, and detailed documentation of phonetic artifacts. A C2 level fluency and formal auditory training are required for success.

Qualifications

  • 2+ years of dedicated international voice process experience in a Captive / In-House GCC.
  • C2 near-native English proficiency with deep conversational nuance mastery.
  • Formally trained auditory ear for vocal inflections and articulation artifacts.

Responsibilities

  • S2S & TTS naturalness scoring across intelligibility and cadence.
  • Benchmark paralinguistic attributes like empathy, tone inflection, and urgency.
  • Conduct blinded pairwise audio preferences with detailed rationales; verify multi-modal synchronization.
  • Document timestamped acoustic anomalies and artifacts during evaluation.

Skills

C2 English proficiency
Auditory acuity
Voice process experience
Multilingual ability

Education

Linguistics / Phonetics background

Tools

Acoustic annotation tools

Job description

ROLE OVERVIEW & OBJECTIVE

We are looking for exceptional Speech & Audio AI Evaluation Specialists with genuine in-house Global Capability Centre (GCC) / Captive international voice experience to evaluate state-of-the-art Speech-to-Speech (S2S), Text-to-Speech (TTS), and real-time conversational voice agents. Unlike traditional transcription or BPO roles, this specialist position demands a highly trained auditory ear to evaluate prosody, cadence, phonetic accuracy, emotional steering, and paralinguistic nuance across global English dialects. Candidates must possess C2 level near-native fluency to execute rigorous human-preference and synchronization benchmarks.

KEY RESPONSIBILITIES & CORE WORKFLOWS

S2S & TTS Naturalness Scoring: Evaluate live Speech-to-Speech and neural Text-to-Speech outputs across intelligibility, rhythm, and conversational cadence.

  • Paralinguistic & Emotion Steering Evaluation: Benchmark how effectively models express nuanced vocal attributes including empathy, hesitation, tone inflection, irony, and situational urgency.
  • Pairwise Audio Preference Ratings: Conduct blinded, head-to-head auditory preference evaluations between candidate audio completions, providing detailed perceptual rationales. Multi-Modal data validations: Verify multi-modal synchronization, lip-sync alignment, and acoustic scene consistency for video dubbing and avatar-driven speech models.
  • Accent & Dialect Calibration: Apply standardized rubric metrics uniformly across North American, British, Australian, and international English accents without regional bias.
  • Defensible Auditory Documentation: Document timestamped acoustic anomalies, unnatural vocal artifacts, metallic distortion, and hallucinated phonetic segments.
CANDIDATE PROFILE & QUALIFICATIONS
Mandatory Requirements
  • Experience: 2+ years of dedicated international voice process experience exclusively within a Captive / In-House Global Capability Centre (GCC) serving US/global client bases.
  • Language Mastery: C2 near-native English proficiency with absolute mastery over conversational nuances, colloquialisms, idioms, and tonal dynamics.
  • Acoustic Acuity: Formally trained auditory ear for subtle vocal inflections, pitch modulation, cadence shifts, and articulation artifacts.
Preferred Qualifications
  • Linguistics & Phonetics: Academic coursework or practical background in phonetics, phonology, auditory acoustics, or speech-language pathology.
  • Speech QA Background: Prior professional experience in voice quality analytics, acoustic data annotation, or speech synthesis evaluation.
  • Acoustic Acuity: Formally trained auditory ear for subtle vocal inflections, pitch modulation, cadence shifts, and articulation artifacts.
  • Multilingual Ability: Additional fluency in major European, Asian, or Latin American languages to support cross-lingual speech benchmarks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Analyst (Speech and Audio AI evaluation)
Analyst (Speech and Audio AI evaluation)

Innodata Inc. • Dadri

On-site
INR 900,000 - 1,350,000
Speech & Audio AI Evaluation Specialist (International Voice)
Speech & Audio AI Evaluation Specialist (International Voice)

Innodata India • Dadri

On-site
INR 900,000 - 1,500,000
Speech & Audio AI Evaluation Specialist (INT VOICE)
Speech & Audio AI Evaluation Specialist (INT VOICE)

Mindtel • Delhi

On-site
INR 400,000 - 640,000
Fortune 500 deployments
On-site collaboration
Impact-driven role
International Voice Process - Speech Evaluation
International Voice Process - Speech Evaluation

Innodata India • Dadri

On-site
INR 279,000 - 502,000
Applied AI Engineer
Applied AI Engineer

Synth (YC S21) • Bengaluru

On-site
INR 1,500,000 - 2,500,000
QA Lead_XR_Audio Validation
QA Lead_XR_Audio Validation

Quest Global • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram

Vibehackers • Gurugram District

On-site
INR 4,000,000 - 7,000,000
AI Voice Specialist
AI Voice Specialist

Kuku • Mumbai

On-site
INR 600,000 - 1,000,000
AI Engineer (Voice Applications)
AI Engineer (Voice Applications)

Durus Consulting • Chennai District

On-site
INR 400,000 - 750,000
Competitive compensation
Cutting-edge Voice AI projects
Career growth opportunities
Lead Voice AI Engineer
Lead Voice AI Engineer

Simpplr • Gurugram District

Hybrid
INR 4,500,000 - 7,000,000