Research Engineer, Audio and Speech

Decagon

New York (NY)

On-site

USD 200,000 - 400,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision
Life Insurance
Retirement Plan
Parental Leave
Wellness stipend
Office lunches

Job summary

Decagon, a leading conversational AI platform, is seeking a Research Engineer focused on Audio and Speech to build models and agent harnesses powering real-time voice agents, taking them from idea to production.

You will advance multimodal and full-duplex systems that listen, reason, speak, and respond in real time, owning your work end-to-end and making high-impact technical decisions.

Qualifications

  • 2+ years of experience in speech, audio ML, multimodal ML, or production machine learning.
  • Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models.
  • Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
  • Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in ML and signal processing.
  • A track record of taking research ideas from prototype to reliable, measurable production impact.

Responsibilities

  • Design and build next-generation agent harnesses optimized for streaming speech and continuous interaction.
  • Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech.
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers and languages.
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes.
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams.

Skills

Speech/Multi-modal ML
Streaming inference
Python & PyTorch
Production impact

Tools

PyTorch

Job description

About Decagon

Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.

Our technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel.

We’re building a future where customer experiences are being redefined from support tickets and hold music to faster resolutions, richer conversations, and deeper relationships. We’re proud to be backed by world-class investors who share that vision, including a16z, Accel, Bain Capital Ventures, Coatue, and Index Ventures, along with many others.

We’re an in-office company, driven by a shared commitment to excellence and velocity. Our values — Just Get It Done, Invent What Customers Want, Winner’s Mindset, and The Polymath Principle — shape how we work and grow as a team.

About the Team

Read more about the Speech Research Team's work:

  • https://decagon.ai/blog/audio-native-semantic-speaker-change-detection
  • https://decagon.ai/blog/scaling-real-time-tts-inference
  • https://decagon.ai/blog/teaching-flow-matching-tts-with-rl

The Research team develops the model and decision-making stack that powers Decagon’s conversational agents for enterprise support. We research, adapt, and implement state-of-the-art techniques in model training, prompting, orchestration, and evaluation in order to make our agents more accurate, robust, and efficient in real-world deployments.

Our goal is to push the frontier of applied conversational AI: agents that reliably understand nuanced intent, track long context, and take the right actions under uncertainty. We measure success the way customers feel it: higher resolution rates, better user satisfaction, and consistent behavior at scale.

About the Role

As a Research Engineer focused on Audio and Speech, you’ll be responsible for building the models and agent harnesses that power Decagon’s real-time voice agents and taking them all the way from idea to production. Your work will advance multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time.

We’re looking for strong engineers who want to build the next generation of AI voice agents. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.

In this role, you will
  • Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction
  • Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale
Your background looks something like this
  • 2+ years of experience in speech, audio ML, multimodal ML, or production machine learning
  • Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models
  • Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio
  • Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing
  • A track record of taking research ideas from prototype to reliable, measurable production impact
Even better if you have
  • Familiarity with speech-to-speech or full-duplex models
  • Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation
Compensation

$200K – $400K + Offers Equity

Benefits

We proudly offer the following benefits for our full-time employees:

  • Medical, Dental, and Vision benefits for you and your family
  • Life Insurance and Disability Benefits
  • Retirement Plan (e.g., 401K, pension)
  • Parental Leave
  • Fertility and family building benefits through Carrot
  • Monthly stipend to support your wellness, lifestyle, and work-life balance
  • Daily lunches and snacks in the office to keep you at your best
  • Take what you need vacation policy (subject to local requirements; UK employees receive 25 days of statutory leave)

These benefits are described in more detail in Decagon’s policies, may vary by location, and can change at any time according to applicable compensation and benefits plans.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (401K, pension)
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, Vision
Life Insurance
401K retirement plan
+5
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Medical, Dental, and Vision
Life Insurance
Disability Benefits
+6
Research Engineer, Audio and Speech
Research Engineer, Audio and Speech

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Senior Research Engineer, Audio and Speech
Senior Research Engineer, Audio and Speech

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+5
Senior Research Engineer, Audio and Speech
Senior Research Engineer, Audio and Speech

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K)
+6
Staff Research Engineer, Voice + Speech
Staff Research Engineer, Voice + Speech

Decagon • New York (NY)

On-site
USD 200,000 - 400,000
Take what you need vacation policy
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
+4
Senior Research Engineer, Voice + Speech
Senior Research Engineer, Voice + Speech

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Vacation policy
Medical, Dental, and Vision benefits
Life Insurance and Disability Benefits
+4
Engineering Manager, Research
Engineering Manager, Research

Decagon • San Francisco (CA)

On-site
USD 280,000 - 430,000
Medical, Dental, Vision
Retirement Plan
Parental Leave
+4
Engineering Manager, Research
Engineering Manager, Research

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 430,000
Medical, Dental, Vision
Life Insurance
Disability Benefits
+6