Senior Speech AI Engineer – Real-Time Voice

Decagon

New York (NY)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
Parental Leave
Fertility and family building benefits
Monthly wellness stipend
Daily lunches and snacks
Take what you need vacation policy (UK
25 days statutory leave (UK)

Job summary

Decagon, an in‑office AI company in the United States, is seeking a Research Engineer focused on Audio and Speech to build real‑time voice agents and production‑ready systems.

You will own end‑to‑end development from idea to deployment, advancing multimodal and full‑duplex models that listen, reason, speak, and respond in real time.

Strong engineers with 4+ years in speech, Python, and PyTorch will thrive; collaboration with platform and infrastructure teams is essential.

Qualifications

  • 4+ years of experience in speech, audio ML, multimodal ML, or production machine learning.
  • Experience developing or adapting autoregressive, diffusion, flow‑matching, or codec‑based speech models.
  • Hands‑on experience with streaming agent systems, low‑latency inference, production model serving, and evaluation on real‑world audio.
  • Fluency in Python and a modern deep‑learning framework such as PyTorch, with strong foundations in machine learning and signal processing.
  • A track record of taking research ideas from prototype to reliable, measurable production impact.

Responsibilities

  • Design and build next‑generation agent harnesses optimized for streaming speech and turn‑taking.
  • Research and train multimodal and full‑duplex models that understand audio and generate speech.
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers and languages.
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, and outcomes.
  • Optimize end‑to‑end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infra teams.

Skills

Speech ML
Multimodal ML
Streaming inference
Python / PyTorch
Production deployment

Tools

PyTorch

Job description

Decagon, an in‑office AI company in the United States, is seeking a Research Engineer focused on Audio and Speech to build real‑time voice agents and production‑ready systems.

You will own end‑to‑end development from idea to deployment, advancing multimodal and full‑duplex models that listen, reason, speak, and respond in real time.

Strong engineers with 4+ years in speech, Python, and PyTorch will thrive; collaboration with platform and infrastructure teams is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Voice AI Research Engineer - Real-Time Speech
Voice AI Research Engineer - Real-Time Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Senior Speech AI Research Engineer — Real-Time Voice
Senior Speech AI Research Engineer — Real-Time Voice

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Audio & Speech Research Engineer – Real-Time Voice AI
Audio & Speech Research Engineer – Real-Time Voice AI

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
Retirement Plan
+3
Real-Time Voice & Speech AI Research Engineer
Real-Time Voice & Speech AI Research Engineer

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Speech & Audio Research Engineer - Real-Time Voice AI
Speech & Audio Research Engineer - Real-Time Voice AI

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Real-Time Voice AI Research Engineer
Real-Time Voice AI Research Engineer

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision
Life Insurance
Disability Benefits
+6
Lead Voice AI Research Engineer – Real-Time Agents
Lead Voice AI Research Engineer – Real-Time Agents

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Take what you need vacation policy
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
+4
Lead Real-Time Voice Platform Architect
Lead Real-Time Voice Platform Architect

Decagon • United States

Remote
USD 180,000 - 260,000
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
Retirement Plan
+3