Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Tencent in Bellevue, WA is seeking researchers to develop large-scale, native multimodal model systems that jointly support vision, audio, and text for comprehensive perception and understanding of the physical world.
The role focuses on end‑to‑end speech models, representation learning, and cross‑modal alignment with potential collaboration across image and text modalities. Requires advanced degrees and strong background in speech processing.
We are building large‑scale, native multimodal model systems that jointly support vision, audio, and text to enable comprehensive perception and understanding of the physical world.
State(s): US-Washington-Bellevue
The expected base pay range for this position is $122,500.00 – $229,700.00 per year. Actual pay may vary depending on job‑related knowledge, skills, and experience.
Employees may be eligible for a sign‑on payment, relocation package, restricted stock units, medical, dental, vision, life and disability benefits, and participation in the company’s 401(k) plan.
Vacation: up to 15 to 25 days per year (depending on tenure). Holidays: up to 13 days per year. Paid sick leave: up to 10 days per year.
We are an equal‑opportunity employer. We firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. Every employee of Tencent feels supported and inspired to achieve individual and common goals.