Principal Speech Data Linguist

Innodata Inc.

United States

Remote

USD 140,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innodata seeks a principal linguist to own linguistic standards and quality for high-volume transcription and segmentation across languages and domains.

You will define conventions, lead QA lifecycles, and shape human-in-the-loop workflows to guide ASR and related models. Bring IPA fluency, deep phonetics expertise, and experience with Praat and ELAN.

Qualifications

  • Substantial industry experience in transcription, segmentation, and speech-data quality.
  • Bachelor's degree in linguistics or closely related field with phonetics foundation.
  • Fluency in IPA and phonetic transcription; experience with acoustic analysis.

Responsibilities

  • Own linguistic standards for transcription and segmentation across languages and use cases.
  • Define transcription/segmentation standards, style guides, and annotation conventions.
  • Establish quality frameworks: rubrics, error taxonomies, adjudication processes, IA and QA at scale.
  • Manage end-to-end quality lifecycle: pre-processing, checks, post-processing, delivery packaging.
  • Design human-in-the-loop workflows to maximize value as ASR quality improves.
  • Handle challenging cases: accented/dialectal speech, low-resource languages, noisy audio.
  • Collaborate with Speech & Audio Research Scientist to align model objectives with transcription specs.
  • Train and mentor transcribers and reviewers for scalable quality.

Skills

IPA transcription
Phonetics
Acoustic analysis
Pronunciation analysis
Multilingual data handling
Scripting in Python
Adjudication and QA
Annotated data pipelines

Education

Bachelor's in linguistics/phonetics/computational linguistics

Tools

Praat
Montreal Forced Aligner
ELAN
Whisper (and ASR engines)
Inter-annotator agreement tools

Job description

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years.

This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself.

What You’ll Own:
  • You will own the linguistic foundation of Innodata’s segmentation and transcription work across languages, domains, and use cases. Concretely, you will:

  • Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, timestamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions.

  • Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale.

  • Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale.

  • Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can’t do rather than on what they already can.

  • Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions.

  • Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models.

  • Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep quality consistent as the work scales.

  • Represent Innodata’s transcription and segmentation approach to the customers and frontier labs we partner with, and contribute to the methodology and best-practice documentation that make our work legible to their teams.

You’ll Thrive in This Role If You Have:
  • Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality — enough that you have authored standards, not only followed them. This is a principal-level role, and we weight practical depth heavily.

  • A Bachelor’s degree in linguistics, phonetics, or computational linguistics, or a closely related field, is required — with a strong foundation in phonetics, phonology, and sociolinguistics so that IPA, prosody, disfluency, dialect, and register are native concepts. An advanced degree is preferred.

  • A big-picture grasp of how transcription and segmentation choices flow downstream into modeling — how different methods change what speech and content-understanding models learn, and therefore which method fits which modeling objective. You can explain to a model builder why a transcription decision matters.

  • Fluency in phonetic transcription and IPA, plus hands-on experience with acoustic and phonetic analysis of speech — spectrograms, formants, pitch and prosody, and segment boundaries — applied to real, messy speech data at scale (for example in Praat).

  • Deep experience with audio segmentation and its conventions — utterance and turn boundaries, timestamping, and speaker and diarization labeling — across real-world audio.

  • Hands-on fluency with the modern speech stack: Whisper and the commercial ASR engines your partners benchmark against (such as AssemblyAI, Deepgram, Rev, and Speechmatics), forced alignment (for example the Montreal Forced Aligner), and annotation tools such as ELAN.

  • Comfort scripting for speech-data work — Python for batch processing, QA, and metrics such as inter-annotator agreement and WER, plus regular expressions and Praat scripting — enough to work fluently with data and pipelines without needing an engineer for every task.

  • Practical data-management skills across the delivery lifecycle — pre-processing, acceptance-stage quality checks, post-processing, pre-delivery validation, and report creation and packaging — so deliverables leave the door correct, consistent, and well documented.

  • Multilingual capability and hands-on experience with accented, dialectal, and code-switched speech; low-resource languages a strong plus.

  • A point of view on how human-in-the-loop workflows should evolve as models improve — where humans stay in the loop, where they move up to adjudication and standard-setting, and how to measure the difference.

  • Strong written and verbal communication, comfortable working directly with research scientists and interfacing with the customers and frontier labs we partner with.

  • Bonus: responsible-AI considerations for speech, such as bias across accents and dialects and privacy and consent in voice data.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Speech & Audio
Research Scientist, Speech & Audio

Innodata Inc. • United States

On-site
USD 160,000 - 185,000
Staff Conversational Designer (PST)
Staff Conversational Designer (PST)

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 190,000
Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)
Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)

OutcomesAI, Inc. • Boston (MA)

Hybrid
USD 150,000 - 230,000
Senior Program Manager, Data Operations
Senior Program Manager, Data Operations

Deepgram • Northern (KY)

Hybrid
USD 90,000 - 120,000
Senior Machine Learning Engineer, Speech & LLM Training Data
Senior Machine Learning Engineer, Speech & LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 190,000
Strategic Projects Lead — Audio Data
Strategic Projects Lead — Audio Data

Besimple AI • San Mateo (CA)

On-site
USD 90,000 - 130,000
Senior Data Scientist - NLP/LLM Specialist
Senior Data Scientist - NLP/LLM Specialist

Scismic • San Diego (CA)

On-site
USD 155,000 - 240,000
Unlimited PTO
401k program
Month-long sabbatical
+2
AI Quality Analyst - Flexible Hours
AI Quality Analyst - Flexible Hours

Innodata Inc. • United States

Hybrid
USD 17,000 - 24,000
Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

Deepgram • San Francisco (CA), Ann Arbor (MI)

Hybrid
USD 250,000 - 420,000
Staff Product Designer, Conversational AI
Staff Product Designer, Conversational AI

Deepgram • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000