ML Researcher, Speech

DeepRec.ai

San Francisco (CA)

Hybrid

USD 250,000 - 300,000

Full time

4 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare
Dental & Vision
Equity
Remote-friendly (US)

Job summary

DeepRec is seeking a Machine Learning Researcher focused on audio to advance production-grade speech technologies. You will develop TTS, STT, and neural audio codecs, translating research into scalable services.

You will collaborate with engineering and product teams, train large models on large audio datasets, and ensure real-time performance in production. Remote-friendly across the US with hybrid options in SF.

Qualifications

  • Strong Python skills and experience training or running large models.
  • Hands-on experience building or improving TTS, ASR, or neural audio codecs.
  • Track record of original, hands-on technical work.
  • Ability to work autonomously and manage experiments.

Responsibilities

  • Build TTS models that sound natural and expressive.
  • Develop STT systems robust to accents, noise, and phone quality.
  • Research neural audio codecs and efficient compression.
  • Explore LLM-audio understanding and real-time inference.
  • Prepare large audio datasets and scale training across GPUs.

Skills

Python
TTS/ASR
Research experience
Real-time systems
Autonomy

Job description

Machine Learning Researcher, Audio

$250,000 – 300,000+, Equity + Bonus

Remote (US & Europe) / San Francisco, CA (Hybrid preferred)

Full-time / Permanent

DeepRec has partnered with a fast-growing, revenue-generating voice AI company empowering enterprises to build AI phone agents at scale. Recent Series C fundingwith backing from leading Silicon Valley investors, they are building the models and infrastructure that make voice the primary interface between businesses and their customers.

This company has built all of their current models completely in-house, and every model ships to real, paying customers almost immediately. No speculative research track here. If you want your work to hit production within weeks, not years, this is that role.

The Opportunity

The research team are working toward a single, ambitious goal: a fully speech-to-speech conversational AI model that understands and responds like a human, in real time.

You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly.

You'll take ideas from theory through large-scale training to production inference serving millions of calls a day, working closely with engineering and product teams to get your research into real customer environments fast.

What You'll Do
  • BuildTTS models that sound natural, expressive, and human
  • BuildSTT systems that stay accurate with accents, background noise, and messy phone lines
  • Work on neural audio codecs, compressing audio efficiently without losing quality
  • Explore LLM-Audio Understanding
  • Prepare and manage large audio datasets, and design how models are trained on them
  • Run training across many GPUs at once, keeping an eye on cost and speed
  • Run fast, well-designed experiments to test what actually works
  • Make sure models run fast and reliably in production, not just in benchmarks
What You'll Bring
Essential
  • A genuine, self-driven interest in this area of research, shown through your own projects or papers, not just what a past employer asked you to do
  • Hands-on experience building or improving TTS, ASR, speech-to-speech, or neural audio codec systems
  • A track record of original, hands-on technical work, rather than off-the-shelf tools or common tutorial-style projects
  • Strong Python skills and experience training or running large models
  • Comfortableworking autonomously
Desirable
  • Experience getting LLMs to understand or reason about audio - the team's top priority right now
  • Experience with model distillation
  • Experience fine-tuning or training language or speech models with reinforcement learning
  • Published research or open-source work in speech or language AI
  • Background working with real-time speech systems or phone-based systems
  • A PhD is welcome but not required. Strong, independent work matters more than qualifications

We encourage you to apply even if you don't meet every requirement. The right mindset and genuine curiosity matter as much as the resume.

What's In It For You
  • Ground-floor seatin a lean and growing research team with real scope to shape it as it scales
  • Every research output ships to real customers, no long speculative research projects
  • High autonomy to shape your own research direction, tooling, and (for senior hires) the team itself
  • Work across ASR, TTS, neural codecs, and the frontier of LLM-audio understanding & speech-to-speech modelling
  • Well-capitalised, fast-moving environment without big-lab bureaucracy
  • Healthcare, dental, vision, meaningful equity, and every tool you need to succeed
  • Remote-friendly across the US
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director, Text-to-Speech Synthesis Research
Director, Text-to-Speech Synthesis Research

Bot Jobs • Myrtle Point (OR)

Remote
USD 213,000 - 328,000
Machine Learning Intern
Machine Learning Intern

Bland AI • San Francisco (CA)

On-site
USD 40,000 - 65,000
Competitive intern compensation
Mentorship from researchers on front‑f
All the tools you need
+2
Research Engineer, Machine Learning Systems
Research Engineer, Machine Learning Systems

Deepgram • United States

On-site
USD 100,000 - 150,000
Holistic health
Unlimited PTO
Learning stipend
+1
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific Global Solutions, Inc. • Redmond (WA)

On-site
USD 150,000 - 160,000
Competitive compensation
Hybrid/Remote options
GPU infrastructure access
+1
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid

Liquid AI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
100% health premiums
401(k) matching up to 4%
Unlimited PTO
+1
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Machine Learning Research Intern 2027
Machine Learning Research Intern 2027

Phonic • San Francisco (CA)

On-site
USD 45,000 - 65,000
Top-tier compensation
Free meals
Off-site & team events
Member of Technical Staff - Multi-Modal, Audio
Member of Technical Staff - Multi-Modal, Audio

Liquid AI • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
Equity
Health insurance
401(k) matching
+1
Senior Machine Learning Engineer [33222]
Senior Machine Learning Engineer [33222]

Stealth Startup • San Carlos (CA)

On-site
USD 225,000 - 325,000
100% medical, dental, and vision insurance
$70/day DoorDash credit
$200/month wellness reimbursement
+1
ML Engineer: Speech & LLMs
ML Engineer: Speech & LLMs

Knowtex • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Unlimited PTO
+3