Research Scientist, Speech & Audio

Innodata

Ridgefield Park (NJ)

On-site

USD 160,000 - 185,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Innodata, a global data engineering company, seeks a Research Scientist to own the science behind speech and audio data for ASR, TTS, and related models in Generative AI contexts. You will define data specs, evaluation criteria, and experiment-driven methods in collaboration with customers and frontier labs.

You will work with multilingual, accented, low-resource, and code-switched speech, publish benchmarks, and guide annotation teams to ensure data quality.

Qualifications

  • 5+ years in speech or audio ML with practical results.
  • Bachelor’s degree required; advanced degree preferred.
  • Trained and evaluated models with PyTorch fundamentals.
  • Experience with multilingual, accented, code-switched speech.
  • Publications or open-source contributions recognized in the field.

Responsibilities

  • Define data specs for ASR, TTS, and audio models into concrete datasets.
  • Build evaluation methods beyond WER, including robustness and latency.
  • Specify sampling, transcription, and annotation schemas for coverage across languages.
  • Run experiments tying data choices to model improvements.
  • Publish benchmarks and collaborate with customers and labs.

Skills

Speech/Audio ML
PyTorch
Code-switching
Publications/Open-source
Model evaluation

Education

Bachelor's in CS/EE or related
MS or PhD preferred

Tools

ESPnet
NeMo
SpeechBrain
Kaldi
HuggingFace
Forced alignment

Job description

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role

Where models actually differ now is robustness across accents, noise, and code-switching; speaker diarization; the naturalness of generated speech; and latency under streaming. Measuring those honestly, and building the data that trains for them, is gated as much by data and evaluation design as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing speech and audio models, and we are hiring a Research Scientist to own the science behind it.

You will partner directly with the customers and frontier labs building ASR, text-to-speech, speech-to-speech and conversational voice, diarization, and audio-language models, as interested in the data behind them as in the models themselves. Your work is judgment: which conditions and languages a benchmark must cover to be honest, what a transcription convention should be for a given objective, and when an automated metric can be trusted versus when a human ear is required. You will also partner closely with our transcription and linguistics lead, whose standards directly shape what the models learn.

What You’ll Own
  • You will define how Innodata designs, structures, and evaluates audio data for speech and audio models, and you will validate those choices experimentally. Concretely, you will:
  • Translate the requirements of speech and audio models — ASR, text-to-speech and speech generation, speech-to-speech and conversational voice, speaker diarization and verification, audio-language models, and streaming systems — into concrete data specifications: modalities, transcription and annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology that goes past word error rate — semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency — and know when automated metrics hold and when they don't.
  • Decide how existing and incoming audio should be structured, enriched, and sampled for coverage that fits the model objective — across languages, accents, and acoustic conditions (studio, real-world, telephonic), speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech.
  • Partner with the transcription and linguistics lead to turn model objectives into transcription specifications, and to quantify how transcription conventions and quality move ASR and speech-model results.
  • Partner with the audio solutions and engineering team so the audio we collect is built for the model objective: you specify what good data and evaluation require, and they scope programs with customers and capture audio to spec.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations — noisy, accented, and adversarial audio — that surface where speech systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic- and augmented-audio pipeline to turn specifications into operational plans.
You’ll Thrive in This Role If You Have
  • Roughly 5+ years of hands‑on industry experience in speech or audio ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated speech or audio models yourself — ASR, TTS, speech-to-speech, speaker, or audio-language models — with strong PyTorch fundamentals.
  • Fluency in the toolchains and metrics speech work runs on: ESPnet, NeMo, SpeechBrain, or Kaldi, HuggingFace, forced alignment, and WER/CER and the metrics that go beyond them.
  • Hands‑on experience with multilingual, accented, dialectal, low‑resource, or code‑switched speech, and with synthetic or augmented audio (TTS pipelines, noise and room‑response simulation).
  • A way of thinking in datasets: you have built evaluation sets, reasoned about coverage across conditions, and argued about what makes speech data good for a given objective.
  • A track record the field recognizes: first‑author publications or strong open‑source contributions at venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS.
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non‑expert audiences, backed by a rigorous, reproducible approach to experiments and documentation.
  • Bonus: interest or hands‑on experience in responsible‑AI evaluation and red‑teaming — for example spoofing and voice‑cloning robustness, or bias across accents and languages.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy,we do not discriminate on the basis of any protected group status under any applicable law.

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey.Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiringprocess or thereafter. Any information that you do provide will be recorded and maintained in aconfidential file.

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection.As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measurethe effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categoriesis as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service‑connected disability.

A "recently separated veteran" means any veteran during the three‑year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Speech Data Linguist
Principal Speech Data Linguist

Innodata • Ridgefield Park (NJ)

On-site
USD 160,000 - 185,000
Research Scientist, Robotics & World Models
Research Scientist, Robotics & World Models

Innodata • Ridgefield Park (NJ)

Remote
USD 160,000 - 185,000
Generative AI Annotator
Generative AI Annotator

Innodata Inc. • United States

Remote
USD 18,000 - 23,000
Research Scientist, Speech & Audio
Research Scientist, Speech & Audio

Innodata Inc. • United States

On-site
USD 160,000 - 185,000
Account Executive, Enterprise Sales, Federal Practice
Account Executive, Enterprise Sales, Federal Practice

Innodata • Washington

Remote
USD 120,000 - 180,000
Data AI Annotator
Data AI Annotator

Innodata Inc. • Northern (KY)

Hybrid
USD 18,000 - 23,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Greenhouse Software, Inc. • Sandy (UT)

On-site
USD 140,000 - 190,000
Team Lead, Android Core Product - San Mateo, CA, USA
Team Lead, Android Core Product - San Mateo, CA, USA

Speechify • San Mateo (CA)

Remote
USD 140,000 - 200,000
Bonus
Stock options
Technical Training & Quality Manager
Technical Training & Quality Manager

Innodata • Austin (TX)

On-site
USD 145,000 - 175,000
Team Lead, Android Core Product - Raleigh, NC, USA
Team Lead, Android Core Product - Raleigh, NC, USA

Speechify • Raleigh (NC)

Remote
USD 140,000 - 200,000
Stock options
Bonus