Tech Lead – ASR, TTS, Speech LLM, IC, Mentor

Jobtailor

Boston (MA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Boston is seeking a leader to guide end-to-end development of ASR, TTS, and Speech LLMs, from architecture to production deployment.

You will mentor a small team focusing on training, synthetic data, active learning, and inference optimization for healthcare applications, and own the technical roadmap for STT/TTS/Speech LLM.

Qualifications

  • M.S. / Ph.D. in Computer Science, Speech Processing, or related field.
  • 7–10 years of experience in applied ML, at least 3 in speech or multimodal AI.
  • Track record of shipping production ASR/TTS models or inference systems at scale.

Responsibilities

  • Lead the end-to-end technical development of speech models (ASR, TTS, Speech-LLM) — from architecture, training strategy, and evaluation to production deployment.
  • Act as an individual contributor and mentor, guiding a small team working on model training, synthetic data generation, active learning, and inference optimization for healthcare applications.
  • Own the technical roadmap for STT/TTS/Speech LLM model training.

Education

M.S. / Ph.D. in Computer Science
Speech Processing

Tools

Triton Inference Server
Kubernetes
GPU Scaling
Fairseq
ESPnet

Job description

  • Lead the end-to-end technical development of speech models (ASR, TTS, Speech-LLM) — from architecture, training strategy, and evaluation to production deployment.
  • Act as an individual contributor and mentor, guiding a small team working on model training, synthetic data generation, active learning, and inference optimization for healthcare applications.
  • Own the technical roadmap for STT/TTS/Speech LLM model training.
Requirements
  • M.S. / Ph.D. in Computer Science, Speech Processing, or related field.
  • 7–10 years of experience in applied ML, at least 3 in speech or multimodal AI.
  • Track record of shipping production ASR/TTS models or inference systems at scale.
  • Deep expertise in speech models (ASR, TTS, Speech LLM) and training frameworks (PyTorch, NeMo, ESPnet, Fairseq).
  • Proven experience with streaming RNN-T / CTC architectures, LoRA/adapters, and TensorRT optimization.
  • Telephony robustness: Codec augmentation (G.711 μ-law, Opus, packet loss/jitter), AGC/loudness norm, band-limit (300–3400 Hz), far-field/noise simulation.
  • Strong understanding of telephony noise, codecs, and real-world audio variability.
  • Experience in Speaker Diarization, turn detection model, smart voice activity detectionEvaluation: WER/latency curves, Entity-F1 (names/DOB/meds), confidence metrics.
  • TTS : VITS/FastPitch/Glow-TTS/Grad-TTS/StyleTTS2, CosyVoice/NaturalSpeech-3 style transfer, BigVGAN/UnivNet vocoders, zero-shot cloning.
  • Speech LLM: Model development and integration with Voice agent pipeline.
  • Experience deploying models with Triton Inference Server, Kubernetes, and GPU scaling.
  • Hands-on with evaluation metrics (WER, F1 on entities, latency p50/p95).
  • Familiarity with LM biasing, WFST grammars, and context injection.
  • Strong mentorship and code-review discipline.
Core Competencies

Demonstrates deep expertise in developing and deploying speech models, including ASR, TTS, and Speech LLM, with a strong focus on production readiness and performance optimization. Proven ability to mentor teams and guide technical roadmaps in healthcare applications.

Highest-signal resume keywords
  • Speech Model Development
  • Production ASR/TTS Model Shipping
  • Applied Machine Learning Expertise
  • Experience with PyTorch and NeMo
  • Telephony Robustness and Noise Handling
ATS Optimization Keywords
Hard Skills
  • ASR Model Development
  • TTS Model Development
  • Speech LLM Integration
  • Streaming RNN-T Architectures
  • TensorRT Optimization
  • Evaluation Metrics (WER, F1)
  • Speaker Diarization
  • Turn Detection Model
  • Smart Voice Activity Detection
  • Codec Augmentation
Soft Skills
  • Mentorship
  • Code Review Discipline
Certifications & Qualifications
  • M.S. / Ph.D. in Computer Science
  • Speech Processing
Industry Keywords
  • Healthcare Applications
  • Multimodal AI
  • Telephony Noise
  • Real-World Audio Variability
  • Context Injection
Tools & Technologies
  • Triton Inference Server
  • Kubernetes
  • GPU Scaling
  • Fairseq
  • ESPnet
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)
Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)

OutcomesAI • Boston (MA)

On-site
USD 180,000 - 260,000
Software Engineer, AI Research – Prototyping
Software Engineer, AI Research – Prototyping

Jobtailor • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Staff AI Engineer
Staff AI Engineer

Jobtailor • Menlo Park (CA)

On-site
USD 150,000 - 190,000
Senior Backend Software Engineer
Senior Backend Software Engineer

Jobtailor • California (MO)

On-site
USD 120,000 - 180,000
ML Engineer: Speech & LLMs
ML Engineer: Speech & LLMs

Knowtex • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Equity
Unlimited PTO
+3
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Edge AI Researcher (Speech & Audio Models)
Edge AI Researcher (Speech & Audio Models)

Huxley • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead Specialist, AI Scientist
Lead Specialist, AI Scientist

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Applied AI Scientist, Senior/Staff
Applied AI Scientist, Senior/Staff

Jobtailor • United States

On-site
USD 120,000 - 180,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000