Lead Speech & Audio Foundation Models Scientist (Remote)

Intelix.AI

San Francisco (CA)

Hybrid

USD 270,000 - 500,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Visa transfer support

Job summary

Intelix.AI is seeking a Principal Research Scientist to lead end-to-end development of foundation models for speech and audio, including text-to-speech and voice cloning. You will frame questions, run experiments, and ship results into production with measurable impact.

Role requires PhD in ML/NLP or equivalent practical experience, and a track record of producing deployable research. Relocation and visa transfer support are available for this US‑based, on-site friendly position with remote

Qualifications

  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Publications or open-source contributions in ML or speech.
  • Experience shipping models to production with measurable impact.

Responsibilities

  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis.
  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling.
  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count.
  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Public work: side projects, open-source, technical write-ups.

Job description

Intelix.AI is seeking a Principal Research Scientist to lead end-to-end development of foundation models for speech and audio, including text-to-speech and voice cloning. You will frame questions, run experiments, and ship results into production with measurable impact.

Role requires PhD in ML/NLP or equivalent practical experience, and a track record of producing deployable research. Relocation and visa transfer support are available for this US‑based, on-site friendly position with remote

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Research Scientist Speech Voice Foundation Models
Principal Research Scientist Speech Voice Foundation Models

Intelix.AI • San Francisco (CA)

Hybrid
USD 270,000 - 500,000
Relocation assistance
Visa transfer support
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Research Scientist - Audio [33340]
Research Scientist - Audio [33340]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Audio Scientist: LALMs & S2S (Remote)
Senior AI Audio Scientist: LALMs & S2S (Remote)

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Senior ML Research Engineer - Speech & Audio (End-to-End)
Senior ML Research Engineer - Speech & Audio (End-to-End)

AI Talent Now • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Scientist - Audio [33341]
Research Scientist - Audio [33341]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Research Scientist - Audio [33363]
Research Scientist - Audio [33363]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Speech LLM Engineer — Build Audio AI (Equity, Hybrid)
Speech LLM Engineer — Build Audio AI (Equity, Hybrid)

Plaud • San Francisco (CA)

Hybrid
USD 195,000 - 365,000
Founding team initiative
Base salary + bonus + Equity
Healthcare benefits
+5
Speech ML Scientist - TTS & Voice AI (Remote)
Speech ML Scientist - TTS & Voice AI (Remote)

Rime • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote-friendly
Visa sponsorship available
Equity upside
+1