Staff ML Engineer: Real-Time Voice Inference Lead

Togetherai

San Francisco (CA)

On-site

USD 220,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Startup equity
Health insurance
Other competitive benefits

Job summary

Togetherai is looking for a Staff ML Engineer based in San Francisco, California to optimize their voice AI platform. In this role, you will influence the model serving layer for voice workloads while working with advanced inference engines.

The ideal candidate should have over 8 years of experience in ML engineering, particularly focused on model serving, with proven expertise in Python and GPU optimization.

Qualifications

  • 8+ years of ML engineering experience, with a focus on model serving.
  • Proficiency in LLM serving engines and GPU optimization.
  • Experience with architectural decisions that influenced platform evolution.

Responsibilities

  • Drive technical strategy for optimizing voice model serving.
  • Lead productionization of voice models at scale.
  • Build a rigorous model evaluation framework.

Skills

ML engineering experience
Model serving expertise
Python proficiency
GPU optimization
System design judgment
Technical leadership
Speech and audio ML knowledge

Education

Bachelor's or Master's in Computer Science or related field

Tools

PyTorch
TensorRT-LLM
CUDA

Job description

Togetherai is looking for a Staff ML Engineer based in San Francisco, California to optimize their voice AI platform. In this role, you will influence the model serving layer for voice workloads while working with advanced inference engines.

The ideal candidate should have over 8 years of experience in ML engineering, particularly focused on model serving, with proven expertise in Python and GPU optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer, Voice AI - Real-Time Inference Lead
Staff ML Engineer, Voice AI - Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Staff ML Engineer — Real-Time Voice Inference Architect
Staff ML Engineer — Real-Time Voice Inference Architect

Together • San Francisco (CA)

On-site
USD 220,000 - 280,000
Senior ML Engineer, Voice AI — Real-Time Inference at Scale
Senior ML Engineer, Voice AI — Real-Time Inference at Scale

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Competitive salary
Startup equity
Health insurance
+1
Senior ML Engineer, Voice AI — Real-Time Inference Lead
Senior ML Engineer, Voice AI — Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Founding Senior ML Engineer, Real-Time Voice AI
Founding Senior ML Engineer, Real-Time Voice AI

Retell Ai • San Francisco (CA)

On-site
USD 143,000 - 215,000
Senior Staff Engineer, Voice AI Infrastructure & ML
Senior Staff Engineer, Voice AI Infrastructure & ML

Connect-AI • San Francisco (CA)

On-site
USD 160,000 - 300,000
Staff Platform Architect — Real-Time Voice AI
Staff Platform Architect — Real-Time Voice AI

Togetherai • San Francisco (CA)

On-site
USD 220,000 - 280,000
Equity
Health insurance
Competitive benefits
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Voice AI Platform Engineer - Real-Time Inference
Voice AI Platform Engineer - Real-Time Inference

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation, including meaningful equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy including company-wide Winter Break
+3
Lead Real-Time Voice AI Architect
Lead Real-Time Voice AI Architect

Inflection AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 400,000 - 550,000
Robust medical, dental and vision with
401k matching
Flexible Time Off
+3