Speech & Audio Research Engineer - Real-Time Voice AI

Decagon

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
Parental Leave
Fertility and family building benefits
Wellness stipend
Daily lunches and snacks
Generous vacation policy

Job summary

Decagon is building a leading conversational AI platform focused on real-time voice agents for enterprise support. As a Research Engineer, you will develop models and harnesses that power Decagon’s voice agents from concept to production, advancing multimodal and full-duplex systems that listen, reason, speak, and respond in real time.

We’re seeking engineers who own their work end-to-end, ship tangible improvements, and make high-impact technical decisions to push the frontier of applied

Qualifications

  • 2+ years in speech, audio ML, multimodal ML, or production ML.
  • Experience with autoregressive, diffusion, flow-matching, or codec-based speech models.
  • Hands-on with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
  • Fluency in Python and PyTorch, with strong foundations in ML and signal processing.
  • Track record of taking research ideas from prototype to reliable, measurable production impact.

Responsibilities

  • Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction
  • Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale

Skills

Speech/Audio ML
Multimodal ML
Production ML
Python

Tools

PyTorch

Job description

Decagon is building a leading conversational AI platform focused on real-time voice agents for enterprise support. As a Research Engineer, you will develop models and harnesses that power Decagon’s voice agents from concept to production, advancing multimodal and full-duplex systems that listen, reason, speak, and respond in real time.

We’re seeking engineers who own their work end-to-end, ship tangible improvements, and make high-impact technical decisions to push the frontier of applied

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Audio & Speech Research Engineer – Real-Time Voice AI
Audio & Speech Research Engineer – Real-Time Voice AI

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
Retirement Plan
+3
Real-Time Voice & Speech AI Research Engineer
Real-Time Voice & Speech AI Research Engineer

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Real-Time Voice AI Research Engineer
Real-Time Voice AI Research Engineer

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision
Life Insurance
Disability Benefits
+6
Senior Speech AI Research Engineer — Real-Time Voice
Senior Speech AI Research Engineer — Real-Time Voice

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Voice AI Research Engineer - Real-Time Speech
Voice AI Research Engineer - Real-Time Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Senior Speech AI Engineer – Real-Time Voice
Senior Speech AI Engineer – Real-Time Voice

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+6
Lead Real-Time Voice Platform Architect
Lead Real-Time Voice Platform Architect

Decagon • United States

Remote
USD 180,000 - 260,000
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision
Life Insurance
Disability Benefits
+6
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
Retirement Plan
+3