Principal Speech Data Linguist

Innodata

Ridgefield Park (NJ)

On-site

USD 160,000 - 185,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Innodata seeks a principal linguist to own linguistic standards for segmentation and transcription across languages and use cases.

You will define standards, build QA frameworks, and guide human-in-the-loop workflows to improve ASR and related models with scalable quality control.

Qualifications

  • Extensive experience in transcription and segmentation quality.
  • Phonetics, IPA, prosody, dialect, and sociolinguistics knowledge.
  • Ability to set standards and write guidelines at scale.
  • Experience with audio segmentation conventions and diarization.

Responsibilities

  • Define transcription and segmentation standards, style guides, and annotation conventions.
  • Establish and run quality frameworks: rubrics, taxonomies, adjudication, and QA at scale.
  • Own end-to-end quality lifecycle for deliverables and reporting.
  • Design human-in-the-loop workflows and determine value-added review steps.
  • Train and mentor transcribers and reviewers; onboard calibration for scale.

Skills

Phonetics
IPA
Acoustic analysis
Praat scripting
Python for QA
Inter-annotator agreement
Multilingual speech
Speech segmentation
Quality frameworks

Education

Bachelor's degree in linguistics / phonetics / computational linguistics

Tools

Praat
Montreal Forced Aligner
ELAN
Whisper / ASR engines

Job description

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years.

This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself.

What You’ll Own:

  • You will own the linguistic foundation of Innodata's segmentation and transcription work across languages, domains, and use cases. Concretely, you will:
  • Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, timestamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions.
  • Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale.
  • Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale.
  • Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can't do rather than on what they already can.
  • Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions.
  • Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models.
  • Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep quality consistent as the work scales.
  • Represent Innodata's transcription and segmentation approach to the customers and frontier labs we partner with, and contribute to the methodology and best-practice documentation that make our work legible to their teams.

You’ll Thrive in This Role If You Have:

  • Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality — enough that you have authored standards, not only followed them. This is a principal-level role, and we weight practical depth heavily.
  • A Bachelor's degree in linguistics, phonetics, or computational linguistics, or a closely related field, is required — with a strong foundation in phonetics, phonology, and sociolinguistics so that IPA, prosody, disfluency, dialect, and register are native concepts. An advanced degree is preferred.
  • A big-picture grasp of how transcription and segmentation choices flow downstream into modeling — how different methods change what speech and content-understanding models learn, and therefore which method fits which modeling objective. You can explain to a model builder why a transcription decision matters.
  • Fluency in phonetic transcription and IPA, plus hands-on experience with acoustic and phonetic analysis of speech — spectrograms, formants, pitch and prosody, and segment boundaries — applied to real, messy speech data at scale (for example in Praat).
  • Deep experience with audio segmentation and its conventions — utterance and turn boundaries, timestamping, and speaker and diarization labeling — across real-world audio.
  • Hands-on fluency with the modern speech stack: Whisper and the commercial ASR engines your partners benchmark against (such as AssemblyAI, Deepgram, Rev, and Speechmatics), forced alignment (for example the Montreal Forced Aligner), and annotation tools such as ELAN.
  • Comfort scripting for speech-data work — Python for batch processing, QA, and metrics such as inter-annotator agreement and WER, plus regular expressions and Praat scripting — enough to work fluently with data and pipelines without needing an engineer for every task.
  • Practical data-management skills across the delivery lifecycle — pre-processing, acceptance-stage quality checks, post-processing, pre-delivery validation, and report creation and packaging — so deliverables leave the door correct, consistent, and well documented.
  • Multilingual capability and hands-on experience with accented, dialectal, and code-switched speech; low-resource languages a strong plus.
  • A point of view on how human-in-the-loop workflows should evolve as models improve — where humans stay in the loop, where they move up to adjudication and standard-setting, and how to measure the difference.
  • Strong written and verbal communication, comfortable working directly with research scientists and interfacing with the customers and frontier labs we partner with.
  • Bonus: responsible-AI considerations for speech, such as bias across accents and dialects and privacy and consent in voice data.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey.Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiringprocess or thereafter. Any information that you do provide will be recorded and maintained in aconfidential file.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy,we do not discriminate on the basis of any protected group status under any applicable law.

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection.As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measurethe effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categoriesis as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Speech & Audio
Research Scientist, Speech & Audio

Innodata • Ridgefield Park (NJ)

On-site
USD 160,000 - 185,000
Generative AI Annotator
Generative AI Annotator

Innodata Inc. • United States

Remote
USD 18,000 - 23,000
Research Scientist, Robotics & World Models
Research Scientist, Robotics & World Models

Innodata • Ridgefield Park (NJ)

Remote
USD 160,000 - 185,000
Account Executive, Enterprise Sales, Federal Practice
Account Executive, Enterprise Sales, Federal Practice

Innodata • Washington

Remote
USD 120,000 - 180,000
Data AI Annotator
Data AI Annotator

Innodata Inc. • Northern (KY)

Hybrid
USD 18,000 - 23,000
Technical Training & Quality Manager
Technical Training & Quality Manager

Innodata • Austin (TX)

On-site
USD 145,000 - 175,000
Staff+ Software Engineer, People Products
Staff+ Software Engineer, People Products

Anthropic • United States

Hybrid
USD 405,000 - 485,000
Data Annotator
Data Annotator

Innodata • Washington

On-site
USD 42,000 - 63,000
Team Lead, Android Core Product - Mountain View, CA, USA
Team Lead, Android Core Product - Mountain View, CA, USA

Speechify • Mountain View (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Team Lead, Android Core Product - San Mateo, CA, USA
Team Lead, Android Core Product - San Mateo, CA, USA

Speechify • San Mateo (CA)

Remote
USD 140,000 - 200,000
Bonus
Stock options