Research Scientist – Speech and Audio Understanding (Large Models & Multimodal Systems)

Lightspeed Studios

Bellevue (WA)

On-site

USD 123,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Sign-on bonus
Relocation package
RSUs

Job summary

Tencent in Bellevue, WA is seeking researchers to develop large-scale, native multimodal model systems that jointly support vision, audio, and text for comprehensive perception and understanding of the physical world.

The role focuses on end‑to‑end speech models, representation learning, and cross‑modal alignment with potential collaboration across image and text modalities. Requires advanced degrees and strong background in speech processing.

Qualifications

  • Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Linguistics, or a related field; or Master’s degree with several years of relevant experience.
  • Solid understanding of speech and audio signal processing, acoustic modeling, language modeling, and large model architectures.
  • Proficient in core speech system pipelines such as ASR, TTS, or speech translation; multilingual/multitask/end-to-end experience is a plus.

Responsibilities

  • Develop general‑purpose, end‑to‑end large speech models across multilingual ASR, translation, synthesis, and understanding.
  • Advance research on speech representation learning and encoder/decoder architectures for unified acoustic representations.
  • Explore alignment and fusion between audio/speech and other modalities for multimodal training/inference.
  • Build and maintain high‑quality multimodal speech datasets with annotation and data synthesis.

Skills

Speech processing
Acoustic modeling
Language modeling
Transformer architectures
Multimodal training
Distributed systems
PyTorch
TensorFlow

Education

Ph.D. or Master’s degree in CS/EE/AI/Linguistics
Relevant experience after degree

Tools

PyTorch
TensorFlow

Job description

What the Role Entails

We are building large‑scale, native multimodal model systems that jointly support vision, audio, and text to enable comprehensive perception and understanding of the physical world.

Job Responsibilities
  • Develop general‑purpose, end‑to‑end large speech models covering multilingual automatic speech recognition (ASR), speech translation, speech synthesis, paralinguistic understanding, and general audio understanding.
  • Advance research on speech representation learning and encoder/decoder architectures to build unified acoustic representations for multi‑task and multimodal applications.
  • Explore representation alignment and fusion mechanisms between audio/speech and other modalities in large multimodal models, enabling joint modeling with image and text.
  • Build and maintain high‑quality multimodal speech datasets, including automatic annotation and data synthesis technologies.
Who We Look For
  • Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Linguistics, or a related field; or Master’s degree with several years of relevant experience.
  • Solid understanding of speech and audio signal processing, acoustic modeling, language modeling, and large model architectures.
  • Proficient in one or more core speech system development pipelines such as ASR, TTS, or speech translation; experience with multilingual, multitask, or end‑to‑end systems is a plus.
  • Experience with large‑scale training and distributed systems is a plus.
  • Familiar with Transformer‑based architectures and their applications in speech and multimodal training/inference.
Preferred Expertise
  • Speech representation pretraining (e.g., HuBERT, Wav2Vec, Whisper).
  • Multimodal alignment and cross‑modal modeling (e.g., audio‑visual‑text).
  • Experience driving state‑of‑the‑art (SOTA) performance on audio understanding tasks with large models.
  • Proficient in deep learning frameworks such as PyTorch or TensorFlow.
Location

State(s): US-Washington-Bellevue

Compensation

The expected base pay range for this position is $122,500.00 – $229,700.00 per year. Actual pay may vary depending on job‑related knowledge, skills, and experience.

Benefits

Employees may be eligible for a sign‑on payment, relocation package, restricted stock units, medical, dental, vision, life and disability benefits, and participation in the company’s 401(k) plan.

Vacation: up to 15 to 25 days per year (depending on tenure). Holidays: up to 13 days per year. Paid sick leave: up to 10 days per year.

Equal Employment Opportunity

We are an equal‑opportunity employer. We firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. Every employee of Tencent feels supported and inspired to achieve individual and common goals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Internship- Multimodal LLM (Speech/Music/Audio/Vision/Language)
Research Internship- Multimodal LLM (Speech/Music/Audio/Vision/Language)

Lightspeed Studios • Bellevue (WA)

On-site
USD 80,000 - 120,000
Company-sponsored medical plan
Paid sick leave
Paid holidays
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific Global Solutions, Inc. • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Competitive compensation
Hybrid/Remote options
GPU infrastructure access
+1
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Equity
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Hunyuan Multimodal Algorithm Researcher (Omni-Modal)
Hunyuan Multimodal Algorithm Researcher (Omni-Modal)

Lightspeed Studios • Bellevue (WA)

On-site
Sign-on payment
Relocation package
Restricted stock units
+5
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Principal Applied Scientist, Real-Time Conversational AI , AGI
Principal Applied Scientist, Real-Time Conversational AI , AGI

Amazon Science • Sunnyvale (CA)

On-site
USD 229,000 - 309,000
Health insurance
401(k) matching
Paid time off
+1
Senior Applied Scientist, Real-Time Conversational AI , AGI
Senior Applied Scientist, Real-Time Conversational AI , AGI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Hunyuan Multimodal Algorithm Researcher (Omni-Modal)
Hunyuan Multimodal Algorithm Researcher (Omni-Modal)

Tencent • Palo Alto (CA)

On-site
USD 134,000 - 254,000
Medical, dental, and vision coverage
Relocation package
401(k) plan
+2