- Lead the end-to-end technical development of speech models (ASR, TTS, Speech-LLM) — from architecture, training strategy, and evaluation to production deployment.
- Act as an individual contributor and mentor, guiding a small team working on model training, synthetic data generation, active learning, and inference optimization for healthcare applications.
- Own the technical roadmap for STT/TTS/Speech LLM model training.
Requirements
- M.S. / Ph.D. in Computer Science, Speech Processing, or related field.
- 7–10 years of experience in applied ML, at least 3 in speech or multimodal AI.
- Track record of shipping production ASR/TTS models or inference systems at scale.
- Deep expertise in speech models (ASR, TTS, Speech LLM) and training frameworks (PyTorch, NeMo, ESPnet, Fairseq).
- Proven experience with streaming RNN-T / CTC architectures, LoRA/adapters, and TensorRT optimization.
- Telephony robustness: Codec augmentation (G.711 μ-law, Opus, packet loss/jitter), AGC/loudness norm, band-limit (300–3400 Hz), far-field/noise simulation.
- Strong understanding of telephony noise, codecs, and real-world audio variability.
- Experience in Speaker Diarization, turn detection model, smart voice activity detectionEvaluation: WER/latency curves, Entity-F1 (names/DOB/meds), confidence metrics.
- TTS : VITS/FastPitch/Glow-TTS/Grad-TTS/StyleTTS2, CosyVoice/NaturalSpeech-3 style transfer, BigVGAN/UnivNet vocoders, zero-shot cloning.
- Speech LLM: Model development and integration with Voice agent pipeline.
- Experience deploying models with Triton Inference Server, Kubernetes, and GPU scaling.
- Hands-on with evaluation metrics (WER, F1 on entities, latency p50/p95).
- Familiarity with LM biasing, WFST grammars, and context injection.
- Strong mentorship and code-review discipline.
Core Competencies
Demonstrates deep expertise in developing and deploying speech models, including ASR, TTS, and Speech LLM, with a strong focus on production readiness and performance optimization. Proven ability to mentor teams and guide technical roadmaps in healthcare applications.
Highest-signal resume keywords
- Speech Model Development
- Production ASR/TTS Model Shipping
- Applied Machine Learning Expertise
- Experience with PyTorch and NeMo
- Telephony Robustness and Noise Handling
ATS Optimization Keywords
Hard Skills
- ASR Model Development
- TTS Model Development
- Speech LLM Integration
- Streaming RNN-T Architectures
- TensorRT Optimization
- Evaluation Metrics (WER, F1)
- Speaker Diarization
- Turn Detection Model
- Smart Voice Activity Detection
- Codec Augmentation
Soft Skills
- Mentorship
- Code Review Discipline
Certifications & Qualifications
- M.S. / Ph.D. in Computer Science
- Speech Processing
Industry Keywords
- Healthcare Applications
- Multimodal AI
- Telephony Noise
- Real-World Audio Variability
- Context Injection
Tools & Technologies
- Triton Inference Server
- Kubernetes
- GPU Scaling
- Fairseq
- ESPnet