Senior Speech AI Research Engineer — Real-Time Voice

Decagon

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
Parental Leave
Fertility and family building benefits
Wellness stipend
Daily lunches and snacks
Generous vacation policy

Job summary

Decagon in San Francisco is seeking a Research Engineer focused on Audio and Speech to build models and agent harnesses powering real-time voice agents from idea to production. You will advance multimodal and full-duplex systems that listen, reason, speak, and respond naturally in real time.

We seek engineers who own their work end-to-end, ship tangible improvements, and make high-impact technical decisions. Strong background in speech, ML, and production systems is required.

Qualifications

  • 4+ years of experience in speech, audio ML, multimodal ML, or production machine learning.
  • Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models.
  • Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
  • Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing.
  • A track record of taking research ideas from prototype to reliable, measurable production impact.

Responsibilities

  • Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction.
  • Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech.
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages.
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes.
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale.

Skills

Speech ML
Multimodal ML
Production ML
Python
PyTorch
Signal processing

Tools

Streaming
Low-latency
Model serving

Job description

Decagon in San Francisco is seeking a Research Engineer focused on Audio and Speech to build models and agent harnesses powering real-time voice agents from idea to production. You will advance multimodal and full-duplex systems that listen, reason, speak, and respond naturally in real time.

We seek engineers who own their work end-to-end, ship tangible improvements, and make high-impact technical decisions. Strong background in speech, ML, and production systems is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Speech AI Engineer – Real-Time Voice
Senior Speech AI Engineer – Real-Time Voice

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+6
Real-Time Voice & Speech AI Research Engineer
Real-Time Voice & Speech AI Research Engineer

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Speech & Audio Research Engineer - Real-Time Voice AI
Speech & Audio Research Engineer - Real-Time Voice AI

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Voice AI Research Engineer - Real-Time Speech
Voice AI Research Engineer - Real-Time Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Lead Voice AI Research Engineer (Production)
Lead Voice AI Research Engineer (Production)

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Vacation policy
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
+4
Lead Voice AI Research Engineer – Real-Time Agents
Lead Voice AI Research Engineer – Real-Time Agents

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Take what you need vacation policy
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
+4
Lead Real-Time Voice Platform Architect
Lead Real-Time Voice Platform Architect

Decagon • United States

Remote
USD 180,000 - 260,000
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5