ML Researcher, Speech

DeepRec.ai

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Healthcare
Dental
Vision

Job summary

DeepRec.ai is a fast-growing voice AI company in SF, delivering production-ready AI phone agents. This full-time role focuses on building end-to-end speech technologies across TTS, STT, and neural codecs, with a goal of real-time, human-like interactions for enterprise customers.

You will push theory to production, train on massive audio datasets, and collaborate with engineering and product teams to ship capabilities to customers quickly. A PhD is welcome but not required.

Qualifications

  • PhD preferred but not required; strong publication record or open-source work.
  • Hands-on experience building or improving TTS, ASR, speech-to-speech, or neural audio codec systems.
  • Strong Python skills; experience training or running large models.
  • Experience with LLMs to understand or reason about audio; reinforcement learning knowledge a plus.
  • Background in real-time speech systems or phone-based systems.

Responsibilities

  • Build TTS models that sound natural, expressive, and human.
  • Build STT systems robust to accents, noise, and phone line issues.
  • Develop neural audio codecs for efficient compression without quality loss.
  • Explore LLM-Audio Understanding.
  • Prepare and manage large audio datasets and design training on them.
  • Run training across multi-GPU setups with cost and speed in mind.
  • Run fast, well-designed experiments to test what works.
  • Ensure models run in production with reliability, not just benchmarks.

Skills

Python
TTS/ASR expertise
Speech AI research
Large model training
Autonomy
RLHF / RL training
Publication/open-source

Education

PhD welcome but not required

Tools

PyTorch
TensorFlow
GPU clusters

Job description

Full-time / Permanent

DeepRec has partnered with a fast-growing, revenue-generating voice AI company empowering enterprises to build AI phone agents at scale. Recent Series C funding with backing from leading Silicon Valley investors, they are building the models and infrastructure that make voice the primary interface between businesses and their customers.

This company has built all of their current models completely in-house, and every model ships to real, paying customers almost immediately. No speculative research track here. If you want your work to hit production within weeks, not years, this is that role.

The Opportunity

The research team are working toward a single, ambitious goal: a fully speech-to-speech conversational AI model that understands and responds like a human, in real time. You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly.

You'll take ideas from theory through large-scale training to production inference serving millions of calls a day, working closely with engineering and product teams to get your research into real customer environments fast.

What You'll Do

  • Build TTS models that sound natural, expressive, and human
  • Build STT systems that stay accurate with accents, background noise, and messy phone lines
  • Work on neural audio codecs, compressing audio efficiently without losing quality
  • Explore LLM-Audio Understanding
  • Prepare and manage large audio datasets, and design how models are trained on them
  • Run training across many GPUs at once, keeping an eye on cost and speed
  • Run fast, well-designed experiments to test what actually works
  • Make sure models run fast and reliably in production, not just in benchmarks

What You'll Bring

  • A genuine, self-driven interest in this area of research, shown through your own projects or papers, not just what a past employer asked you to do
  • Hands-on experience building or improving TTS, ASR, speech-to-speech, or neural audio codec systems
  • A track record of original, hands-on technical work, rather than off-the-shelf tools or common tutorial-style projects
  • Strong Python skills and experience training or running large models
  • Comfortable working autonomously
  • Experience getting LLMs to understand or reason about audio - the team's top priority right now
  • Experience with model distillation
  • Experience fine-tuning or training language or speech models with reinforcement learning
  • Published research or open-source work in speech or language AI
  • Background working with real-time speech systems or phone-based systems
  • A PhD is welcome but not required. Strong, independent work matters more than qualifications

We encourage you to apply even if you don't meet every requirement. The right mindset and genuine curiosity matter as much as the resume.

What's In It For You

  • Ground-floor seat in a lean and growing research team with real scope to shape it as it scales
  • Every research output ships to real customers, no long speculative research projects
  • High autonomy to shape your own research direction, tooling, and (for senior hires) the team itself
  • Work across ASR, TTS, neural codecs, and the frontier of LLM-audio understanding & speech-to-speech modelling
  • Well-capitalised, fast-moving environment without big-lab bureaucracy
  • Healthcare, dental, vision, meaningful equity, and every tool you need to succeed
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Audio [33341]
Research Scientist - Audio [33341]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Remote Voice AI Model Research Engineer
Remote Voice AI Model Research Engineer

Deepslate • Germany (OH)

Remote
USD 90,000 - 150,000
True R&D Freedom
Competitive Compensation
Serious Compute
+1
Research Scientist - Audio [33363]
Research Scientist - Audio [33363]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Research Intern 2027
Machine Learning Research Intern 2027

Mixpeek • San Francisco (CA), Northern (KY)

Hybrid
USD 34,000 - 48,000
Competitive internship pay
SF office in-person work
Meals provided in the office
Machine Learning Research Intern 2027
Machine Learning Research Intern 2027

Phonic • San Francisco (CA)

On-site
USD 45,000 - 65,000
Top-tier compensation
Free meals
Off-site & team events
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Research, Audio Expertise
Research, Audio Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Model Research Engineer (m/f/d)
Model Research Engineer (m/f/d)

Deepslate • Germany (OH)

Remote
USD 90,000 - 150,000
True R&D Freedom
Competitive Compensation
Serious Compute
+1
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
ML Engineer: Speech & LLMs
ML Engineer: Speech & LLMs

Knowtex • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Equity
Unlimited PTO
+3