Stand out for this role — generate a tailored resume and cover letter in about a minute.
DeepRec is seeking a Machine Learning Researcher focused on audio to advance production-grade speech technologies. You will develop TTS, STT, and neural audio codecs, translating research into scalable services.
You will collaborate with engineering and product teams, train large models on large audio datasets, and ensure real-time performance in production. Remote-friendly across the US with hybrid options in SF.
$250,000 – 300,000+, Equity + Bonus
Remote (US & Europe) / San Francisco, CA (Hybrid preferred)
Full-time / Permanent
DeepRec has partnered with a fast-growing, revenue-generating voice AI company empowering enterprises to build AI phone agents at scale. Recent Series C fundingwith backing from leading Silicon Valley investors, they are building the models and infrastructure that make voice the primary interface between businesses and their customers.
This company has built all of their current models completely in-house, and every model ships to real, paying customers almost immediately. No speculative research track here. If you want your work to hit production within weeks, not years, this is that role.
The research team are working toward a single, ambitious goal: a fully speech-to-speech conversational AI model that understands and responds like a human, in real time.
You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly.
You'll take ideas from theory through large-scale training to production inference serving millions of calls a day, working closely with engineering and product teams to get your research into real customer environments fast.
We encourage you to apply even if you don't meet every requirement. The right mindset and genuine curiosity matter as much as the resume.