Staff ML Engineer — Real-Time Voice Inference Architect

Together

San Francisco (CA)

On-site

USD 220,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads.

You'll work with state-of-the-art accelerators and collaborate with model partners to bring models to production on Together's platform. This is a foundational role on a small, high-impact team.

Qualifications

  • 8+ years of ML engineering experience focused on model serving or ML infrastructure.
  • Deep, practical expertise in LLM serving engines and production-scale inference.
  • Expert-level Python and PyTorch with GPU optimization fundamentals.
  • Strong system design judgment and ability to guide platform evolution.
  • Autonomy, strong prioritization, and leadership on complex projects.
  • Product intuition for developer tooling and real-time voice applications.
  • Experience with streaming audio, latency budgets, and production ML stacks.

Responsibilities

  • Own the voice inference roadmap end-to-end for STT, TTS, and speech-to-speech.
  • Architect and implement high-performance inference with low latency and high throughput.
  • Lead productionization of voice models at scale with serverless and dedicated endpoints.
  • Build an evaluation platform covering accuracy, naturalness, latency, and provenance.
  • Shape architecture to support next-gen models and encoder/decoder paradigms.
  • Drive partner integrations with model providers and ensure ongoing performance.
  • Diagnose and resolve hard performance issues with measurable shipped improvements.
  • Collaborate with platform engineering to raise the bar for latency and reliability.

Skills

Python
PyTorch
GPU optimization
System design
Leadership
Speech & audio ML
Developer tooling
Autonomous work

Education

Bachelor's or Master's in CS/EE

Tools

vLLM
SGLang
TensorRT-LLM
CUDA

Job description

Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads.

You'll work with state-of-the-art accelerators and collaborate with model partners to bring models to production on Together's platform. This is a foundational role on a small, high-impact team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer, Voice AI - Real-Time Inference Lead
Staff ML Engineer, Voice AI - Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Senior Machine Learning Engineer, Voice AI
Senior Machine Learning Engineer, Voice AI

Together • San Francisco (CA)

On-site
USD 200,000 - 260,000
Equity
Health insurance
Competitive benefits
Senior Machine Learning Engineer, Voice AI
Senior Machine Learning Engineer, Voice AI

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Competitive salary
Startup equity
Health insurance
+1
Senior Machine Learning Engineer, Voice AI
Senior Machine Learning Engineer, Voice AI

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Competitive salary
Startup equity
Health insurance
+1
Staff Machine Learning Engineer, Voice AI
Staff Machine Learning Engineer, Voice AI

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Staff Machine Learning Engineer, Voice AI
Staff Machine Learning Engineer, Voice AI

Together • San Francisco (CA)

On-site
USD 220,000 - 280,000
Senior ML Engineer, Voice AI — Real-Time Inference Lead
Senior ML Engineer, Voice AI — Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Senior ML Engineer, Voice AI — Real-Time Inference at Scale
Senior ML Engineer, Voice AI — Real-Time Inference at Scale

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Competitive salary
Startup equity
Health insurance
+1
Lead Real-Time Voice AI Architect
Lead Real-Time Voice AI Architect

Inflection AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 400,000 - 550,000
Robust medical, dental and vision with
401k matching
Flexible Time Off
+3
Remote Audio Inference Engineer — Fast ML Serving
Remote Audio Inference Engineer — Fast ML Serving

Visa Hunt • New York (NY)

Hybrid
USD 140,000 - 190,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+6