Senior ML Engineer, Voice AI — Real-Time Inference Lead
Together AI
San Francisco (CA)
On-site
USD 200,000 - 260,000
Full time
14 days+
Application generator
Stand out for this role — generate a tailored resume and cover letter in about a minute.
Get past ATS filters
Benefits offered by this job
Competitive salary
Startup equity
Health insurance
Other competitive benefits
Job summary
An AI infrastructure company is looking for a Senior ML Engineer to optimize performance for voice models such as speech-to-text and text-to-speech. You will work on productionizing models on advanced serving engines, focusing on GPU utilization and inference optimization. With over 5 years of experience required, this role offers competitive compensation in San Francisco, California, and invites you to be part of a visionary team shaping the future of voice applications.
Qualifications
5+ years of experience in ML engineering, focusing on model serving and inference optimization.
Hands-on experience with LLM serving engines and optimization.
Strong proficiency in Python and PyTorch; experience with GPU profiling.
Responsibilities
Optimize inference performance for voice models including STT and TTS.
Productionize voice models on serverless and dedicated endpoints.
Build and maintain a voice model evaluation framework.
Skills
ML engineering
Model serving
Inference optimization
Python
PyTorch
CUDA
Education
Bachelor's or Master's degree in Computer Science or related field
Tools
TensorRT‑LLM
vLLM
SGLang
Job description
An AI infrastructure company is looking for a Senior ML Engineer to optimize performance for voice models such as speech-to-text and text-to-speech. You will work on productionizing models on advanced serving engines, focusing on GPU utilization and inference optimization. With over 5 years of experience required, this role offers competitive compensation in San Francisco, California, and invites you to be part of a visionary team shaping the future of voice applications.