Senior Researcher, Speech Synthesis & Multimodal LLMs

L&Q OASIS PTE. LTD.

Singapore

Hybrid

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Tencent's Lightspeed Tech Center is seeking a PhD-level researcher to advance speech synthesis and multimodal LLMs for real-time spoken interaction. You will prototype and productionize TTS, voice conversion, and audio generation within online applications, collaborating with cross-functional teams to scale experiments to deployment.

Strong background in speech processing, diffusion/autoregressive models, and Python-based DL frameworks is essential; publications at top venues are a plus.

Qualifications

  • Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or closely related field.
  • Strong foundation in speech/audio processing and modern generative models.
  • Hands-on experience extending LLMs to speech/audio modalities.
  • Experience with real-time interaction modeling is a plus.

Responsibilities

  • Research and develop advanced speech synthesis and generation algorithms (e.g., TTS, voice conversion, sound/music generation) based on LLMs.
  • Develop and optimize speech and audio synthesis systems for online applications.
  • Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding and generation.
  • Collaborate cross-functionally with research and engineering teams from prototyping to production.

Skills

Python
Deep Learning
Speech Processing
TTS
Multimodal LLM

Education

Ph.D. in CS/EE/Signal Processing

Tools

PyTorch
TensorFlow

Job description

Tencent's Lightspeed Tech Center is seeking a PhD-level researcher to advance speech synthesis and multimodal LLMs for real-time spoken interaction. You will prototype and productionize TTS, voice conversion, and audio generation within online applications, collaborating with cross-functional teams to scale experiments to deployment.

Strong background in speech processing, diffusion/autoregressive models, and Python-based DL frameworks is essential; publications at top venues are a plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Researcher, Speech Synthesis and Multimodal LLM
Senior Researcher, Speech Synthesis and Multimodal LLM

L&Q OASIS PTE. LTD. • Singapore

Hybrid
SGD 120,000 - 180,000
Senior Researcher, Speech Synthesis and Multimodal LLM
Senior Researcher, Speech Synthesis and Multimodal LLM

LightSpeed Studios • Singapore

On-site
SGD 90,000 - 150,000
Speech Synthesis Intern
Speech Synthesis Intern

Tencent • Singapore

On-site
SGD 90,000 - 130,000
Speech Synthesis Research Intern - Gaming Audio & ML
Speech Synthesis Research Intern - Gaming Audio & ML

Lightspeed Studios • Singapore

On-site
SGD 80,000 - 120,000
Senior Multimodal AI Researcher for Games & Agents
Senior Multimodal AI Researcher for Games & Agents

L&Q OASIS PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
Multimodal LLM Research Engineer
Multimodal LLM Research Engineer

Lightspeed Studios • Singapore

On-site
SGD 80,000 - 120,000
Senior Researcher, Multi-Modality
Senior Researcher, Multi-Modality

LightSpeed Studios • Singapore

On-site
SGD 180,000 - 300,000
ML Engineer: Audio Understanding & Speech Tech
ML Engineer: Audio Understanding & Speech Tech

TikTok • Singapore

On-site
SGD 70,000 - 110,000
Senior Researcher, Large Language Models
Senior Researcher, Large Language Models

LightSpeed Studios • Singapore

On-site
SGD 90,000 - 150,000
Senior Researcher, Multi-Modality
Senior Researcher, Multi-Modality

L&Q OASIS PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000