Principal Research Scientist Speech Voice Foundation Models

Intelix.AI

San Francisco (CA)

On-site

USD 270,000 - 500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Visa transfer support

Job summary

Intelix.AI is seeking a Principal Research Scientist to lead end-to-end development of foundation models for speech and audio, including text-to-speech and voice cloning. You will frame questions, run experiments, and ship results into production with measurable impact.

Role requires PhD in ML/NLP or equivalent practical experience, and a track record of producing deployable research. Relocation and visa transfer support are available for this US‑based, on-site friendly position with remote

Qualifications

  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Publications or open-source contributions in ML or speech.
  • Experience shipping models to production with measurable impact.

Responsibilities

  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis.
  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling.
  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count.
  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Public work: side projects, open-source, technical write-ups.

Job description

Principal Research Scientist Speech & Audio Foundation Models

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.


$270,000–$500,000 base plus bonus, equity and benefits (US).


Relocation assistance available. Visa transfer supported


San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.


Permanent, full-time.


A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.


Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.


The role

Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.


Responsibilities


  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.

  • Build and improve voice and speech models across TTS, STT and speech-to-speech.

  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.

  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.

  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.

  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.

  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.

  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.

  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.

  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.

  • PhD in ML or NLP, or equivalent practical experience you can point to.

  • Public work: side projects, open-source, technical write-ups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 100,000 - 130,000
Mountain View, California, USA Staff / Principal Research Scientist - USA
Mountain View, California, USA Staff / Principal Research Scientist - USA

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Bonus + equity
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
ML Researcher, Speech
ML Researcher, Speech

DeepRec.ai • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Healthcare
Dental
+1
Research Scientist
Research Scientist

Phonic, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Free meals
Comprehensive healthcare
+2
Principal Applied Scientist, Real-Time Conversational AI , AGI
Principal Applied Scientist, Real-Time Conversational AI , AGI

Amazon Science • Sunnyvale (CA)

On-site
USD 229,000 - 309,000
Health insurance
401(k) matching
Paid time off
+1
Research Engineer
Research Engineer

Phonic, Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Top-tier compensation
Free meals
Comprehensive health, dental, and vision
+2
Lead Speech & Audio Foundation Models Scientist (Remote)
Lead Speech & Audio Foundation Models Scientist (Remote)

Intelix.AI • San Francisco (CA)

Hybrid
USD 270,000 - 500,000
Relocation assistance
Visa transfer support
Senior Research Scientist - Foundation Models (Remote, CH)
Senior Research Scientist - Foundation Models (Remote, CH)

careers.bitkraft.vc - Jobboard • Indiana (PA)

On-site
USD 148,000 - 221,000