Senior Voice AI Engineer

Engg

United States

Remote

USD 86,000 - 106,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Fully remote

Job summary

Engg is seeking a founding engineer to own the real-time voice layer from speech input to spoken output on a fully remote, globally distributed team.

You will build streaming speech-to-text, TTS, and the transport stack, while targeting latency under 800 ms on real calls. This role requires hands-on work across live audio, latency optimization, and evaluation harnesses to guide product decisions.

Qualifications

  • 5+ years building production software, with 2+ years in voice/real-time audio.
  • Experience shipping end-to-end real-time voice pipelines.
  • Strong Python or TypeScript; comfortable in both.
  • Clear written English for asynchronous communication.

Responsibilities

  • Own streaming speech-to-text, LLM turn-taking, text-to-speech, and transport (telephony/WebRTC).
  • Measure and reduce latency, aiming for first audio under 800 ms on real calls.
  • Address interruptions, barge-in, silence detection, overlapping speech, accents, and mid-sentence changes.
  • Build an evaluation harness from calls, transcripts, and scored turns to detect regressions.
  • Compare voice providers and models via evidence-based testing and act on results.
  • Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
  • Collaborate directly with founders and influence technical decisions in a fast-moving team.

Skills

Python
TypeScript

Tools

LiveKit
Pipecat
Vapi
Twilio Media Streams
Daily
WebSocket

Job description

ABOUT THE ROLE

As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.

WHAT YOU'LL DO
  • Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
  • Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.
  • Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
  • Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
  • Compare voice providers and models through evidence-based testing, and make changes based on results.
  • Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
  • Work directly with founders and make technical decisions in a fast-moving team.
WHAT WE'RE LOOKING FOR
  • At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.
  • Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.
  • Strong Python or TypeScript skills and comfort working in both.
  • Hands‑on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.
  • Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.
  • Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.
  • Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.
COMPENSATION & BENEFITS

Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.

LOCATION

Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Voice AI Engineer
Senior Voice AI Engineer

Clera • United States

Remote
USD 96,000 - 110,000
Senior Voice AI Engineer - Remote, Real-Time Audio Stack
Senior Voice AI Engineer - Remote, Real-Time Audio Stack

Engg • United States

Remote
USD 86,000 - 106,000
Fully remote
Sr. AI Voice Engineer
Sr. AI Voice Engineer

Midway Auto Group • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Senior Software Engineer
Senior Software Engineer

AIM AI • Los Angeles (CA)

On-site
USD 130,000 - 170,000
Senior Member of Technical Staff
Senior Member of Technical Staff

Thomas Talent Network • San Francisco (CA)

On-site
USD 225,000 - 300,000
Full-Stack Software Engineer
Full-Stack Software Engineer

Bot Jobs • Redwood City (CA)

On-site
USD 215,000 - 290,000
Visa sponsorship
AI Engineer, Voice and Realtime
AI Engineer, Voice and Realtime

Obble • United States

Remote
USD 120,000 - 180,000
Founding AI Engineer
Founding AI Engineer

PetsApp • New York (NY)

On-site
USD 200,000 - 250,000
100% employer-paid health insurance
Flexible PTO
Seed-stage equity grant
+2
Senior / Staff Backend Engineer
Senior / Staff Backend Engineer

Hamming • Austin (TX)

On-site
USD 120,000 - 160,000
Flexible work hours
Career development opportunities
Founding Senior Software Engineer (Full-stack)
Founding Senior Software Engineer (Full-stack)

Stealth Startup • California (MO)

Hybrid
USD 215,000 - 290,000
Unlimited meals
Gym reimbursement
Commute reimbursement
+1