Senior Voice AI Engineer

Clera

United States

Remote

USD 96,000 - 110,000

Full time

48 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Clera is seeking a founding engineer to own the real-time voice layer, from speech input through AI reasoning to spoken output. You will help ensure end-to-end latency is minimized in production, delivering natural, responsive voice interactions.

You will build streaming speech-to-text, text-to-speech, and transport layers while collaborating with the founders on architectural decisions in a fast-moving, fully remote team with core overlap 13:00–17:00 UTC.

Qualifications

  • 5+ years building production software.
  • 2+ years shipping voice/real-time audio.
  • Strong Python or TypeScript skills.
  • Experience with audio stacks such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily.
  • Experience debugging audio at frame level (sample rates, codecs, jitter).
  • Experience building LLM evaluation harnesses and optimizing latency.

Responsibilities

  • Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and transport.
  • Measure and reduce latency, targeting first audio under 800 ms on real calls.
  • Address interruptions, barge-in, silence detection, overlapping speech, accents.
  • Build an evaluation harness from recordings to detect regressions.
  • Compare voice providers and models through evidence-based testing.
  • Instrument production systems for turn latency and cost per minute.
  • Work directly with founders and make technical decisions in a fast-moving team.

Skills

Real-time audio
Latency optimization
Python/TypeScript
Speech tech

Tools

LiveKit
Twilio Media Streams
WebRTC
Pipecat
Vapi
WebSocket
Daily

Job description

About the Role

As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.

What You'll Do
  • Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.

  • Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.

  • Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.

  • Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.

  • Compare voice providers and models through evidence-based testing, and make changes based on results.

  • Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.

  • Work directly with founders and make technical decisions in a fast-moving team.

What We're Looking For
  • At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.

  • Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.

  • Strong Python or TypeScript skills and comfort working in both.

  • Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.

  • Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.

  • Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.

  • Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.

Compensation & Benefits

Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.

Location

Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Voice AI Engineer
Senior Voice AI Engineer

Engg • United States

Remote
USD 86,000 - 106,000
Fully remote
Senior Voice AI Engineer - Remote, Real-Time Audio Stack
Senior Voice AI Engineer - Remote, Real-Time Audio Stack

Engg • United States

Remote
USD 86,000 - 106,000
Fully remote
Founding Software Engineer — AI & Voice Systems
Founding Software Engineer — AI & Voice Systems

Roark • United States

Remote
USD 150,000 - 230,000
Sr. AI Voice Engineer
Sr. AI Voice Engineer

Midway Auto Group • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Senior Software Engineer
Senior Software Engineer

AIM AI • Los Angeles (CA)

On-site
USD 130,000 - 170,000
Full-Stack Software Engineer
Full-Stack Software Engineer

Bot Jobs • Redwood City (CA)

On-site
USD 215,000 - 290,000
Visa sponsorship
Senior Member of Technical Staff
Senior Member of Technical Staff

Thomas Talent Network • San Francisco (CA)

On-site
USD 225,000 - 300,000
Founding AI Engineer
Founding AI Engineer

PetsApp • New York (NY)

On-site
USD 200,000 - 250,000
100% employer-paid health insurance
Flexible PTO
Seed-stage equity grant
+2
AI Engineer, Voice and Realtime
AI Engineer, Voice and Realtime

Obble • United States

Remote
USD 120,000 - 180,000
Senior / Staff Backend Engineer
Senior / Staff Backend Engineer

Hamming • Austin (TX)

On-site
USD 120,000 - 160,000
Flexible work hours
Career development opportunities