Director, Text-to-Speech Synthesis Research

Hidden Jobs

United States

Remote

USD 180,000 - 290,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Deepgram is seeking a director-level research leader to own the end-to-end Text-to-Speech program. You will set strategy, drive architectural choices, and stay hands-on with experiments, training runs, and evaluation to move production metrics.

You will mentor researchers and research engineers, build scalable evaluation systems, and partner with product and engineering leadership to ship high-impact TTS capabilities at production scale.

Qualifications

  • Deep expertise in modern TTS, speech generation, or audio generative modeling.
  • Hands-on track record of training and improving large-scale neural models.
  • Ability to set research direction under uncertainty and prioritize experiments.

Responsibilities

  • Own the TTS research and model roadmap, choosing technical bets to pursue.
  • Drive advances across neural audio modeling, prosody, expressiveness, and multilingual generation.
  • Remain deeply technical: review research, design experiments, diagnose model failures.
  • Build evaluation and benchmarking systems with automated metrics and human perceptual assessment.
  • Lead a mix of individual contributors and tech-lead managers, hiring and developing both.
  • Partner with engineering and product leadership on ship-readiness and represent the program externally.

Skills

TTS research leadership
Neural speech modeling
Experiment design
Leading researchers

Tools

PyTorch
TensorFlow
Distributed training

Job description

Role overview

This director-level research position leads the end-to-end Text-to-Speech program, spanning research strategy, technical direction, and the models that reach production. It is a hands-on leadership role that combines setting roadmap-level direction with staying directly engaged in architectures, experiments, training runs, and evaluation. The mandate is to advance the frontier of neural speech generation while converting those gains into deployable systems that move measurable production metrics.

Responsibilities
  • Own the TTS research and model roadmap, choosing which technical bets are worth pursuing and recognizing when an approach should change or be retired
  • Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual generation, voice identity and consistency, data and training strategy, post-training, and inference performance
  • Remain deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and tackle the highest-leverage problems hands-on
  • Build evaluation and benchmarking systems that explain why models improve, pairing automated metrics with human perceptual assessment
  • Lead a mix of individual contributors and tech-lead managers, hiring and developing both, holding a high technical bar, and setting direction across sub-teams
  • Partner with engineering and product leadership on ship-readiness and represent the program internally and externally
Requirements
  • Deep expertise in modern TTS, speech generation, or audio generative modeling, with a hands-on track record of personally training and improving large-scale neural models
  • Command of the contemporary speech-generation stack and the open problems behind naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost
  • Demonstrated ability to set research direction under genuine uncertainty, prioritizing experiments, allocating compute, and ending approaches that aren't working
  • Experience leading researchers and research engineers through technical leaders, developing tech-lead managers and directing sub-teams while remaining technically influential
  • A working style in which AI tools are the default mode of operation, with an earned view of what they still cannot do in speech research
  • Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences
Nice to have
  • TTS or generative-audio models deployed at meaningful production scale
  • Experience building or substantially scaling a high-performing AI research organization
  • Sophisticated evaluation systems for generative speech, including expressive or multilingual generation, voice cloning, or controllable generation
  • Recognized external contributions through publications, open-source work, patents, or invited talks in speech synthesis, neural audio codecs, speech language models, or multimodal models
  • Background in fast-moving startup or research environments that routinely take models from idea to production
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

Apply • Northern (KY)

Hybrid
USD 180,000 - 260,000
Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

Deepgram • Northern (KY)

Hybrid
USD 180,000 - 320,000
Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

AI Chopping Block • California (MO)

Hybrid
USD 180,000 - 280,000
Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

Deepgram • San Francisco (CA), Ann Arbor (MI)

Hybrid
USD 250,000 - 420,000
Director of Research, Text to Speech
Director of Research, Text to Speech

AI Chopping Block • California (MO), Northern (KY)

Hybrid
USD 200,000 - 320,000
Director of Research, Text to Speech
Director of Research, Text to Speech

Apply • Michigan

On-site
USD 180,000 - 240,000
Director of Research, Text to Speech
Director of Research, Text to Speech

Deepgram • San Francisco (CA), Ann Arbor (MI)

On-site
USD 180,000 - 280,000
Director, TTS Research & Production
Director, TTS Research & Production

Apply • Northern (KY)

Hybrid
USD 180,000 - 260,000
Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 100,000 - 130,000
Director, TTS Research & Production Innovation
Director, TTS Research & Production Innovation

AI Chopping Block • California (MO), Northern (KY)

Hybrid
USD 200,000 - 320,000