Research Scientist - Speech

JAM

United States

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading AI lab in the United States is seeking a Research Scientist to develop advanced speech understanding and generation models. This role involves designing and training cutting-edge models that enhance fidelity and efficiency in speech-to-text, text-to-speech, and speech-to-speech systems. Candidates should have significant research experience in TTS, ASR, or speech synthesis, and the ability to progress ideas from experimentation to production. Join a team shaping the future of voice-driven AI interactions.

Qualifications

  • Strong research experience in TTS, ASR, or speech synthesis.
  • Proven ability to take ideas from experimentation to production.
  • Experience designing and training state-of-the-art speech models.

Responsibilities

  • Design and train models for speech systems.
  • Work with large-scale multi-modal datasets.
  • Move ideas from experimentation to production.

Skills

Speech synthesis
Text-to-speech (TTS)
Automatic speech recognition (ASR)

Job description

Overview

A fast-growing AI lab with compute capabilities on par with top big-tech research groups is looking for a Research Scientist to advance the frontier of speech understanding and generation. You\'ll design and train models that push the limits of fidelity, controllability, and efficiency across speech-to-text, text-to-speech, and speech-to-speech systems. The position blends cutting-edge generative modeling with large-scale multi-modal datasets to enable human-like, expressive voice interactions.

If you have a strong research experience in TTS, ASR, or speech synthesis, with a proven ability to take ideas from experimentation to production then this is a chance to shape the next generation of voice-driven AI experiences.

Responsibilities
  • Design and train models that push the limits of fidelity, controllability, and efficiency across speech-to-text, text-to-speech, and speech-to-speech systems.
  • Work with large-scale multi-modal datasets to enable human-like, expressive voice interactions.
  • Move ideas from experimentation to production through rigorous evaluation and iteration.
Qualifications
  • Strong research experience in TTS, ASR, or speech synthesis.
  • Proven ability to take ideas from experimentation to production.
  • Experience designing and training state-of-the-art speech models.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 140,000 - 230,000
Speech AI Research Scientist: TTS/ASR Innovation
Speech AI Research Scientist: TTS/ASR Innovation

JAM • United States

On-site
USD 100,000 - 130,000
Speech AI Research Scientist: TTS, ASR & Voice Generation
Speech AI Research Scientist: TTS, ASR & Voice Generation

JAM • United States

On-site
USD 140,000 - 230,000
Research Scientist - Audio [33340]
Research Scientist - Audio [33340]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Audio [33363]
Research Scientist - Audio [33363]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Audio [33341]
Research Scientist - Audio [33341]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Principal Research Scientist Speech Voice Foundation Models
Principal Research Scientist Speech Voice Foundation Models

Intelix.AI • San Francisco (CA)

Hybrid
USD 270,000 - 500,000
Relocation assistance
Visa transfer support
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Director of Research, Text to Speech
Director of Research, Text to Speech

Deepgram • Myrtle Point (OR)

On-site
USD 180,000 - 260,000
Director of Research, Text to Speech
Director of Research, Text to Speech

AI Chopping Block • California (MO), Northern (KY)

Hybrid
USD 200,000 - 320,000