Research Scientist, Multimodal AI & Speech

ByteDance

San Jose (CA)

On-site

USD 254,000 - 480,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k)
Parental leave
Short/Long-term disability
Life insurance
Wellbeing programs
Paid holidays
Paid sick days
Paid personal time

Job summary

ByteDance Seed is hiring researchers to push the frontiers of artificial general intelligence. The team explores MLLM, GenMedia, AI for Science, and Robotics, with global labs and opportunities in the United States, China, and Singapore.

Successful candidates will join a collaborative, cutting-edge environment. As a PhD holder, you will contribute to foundational models, speech and multimodal research, and work with world-class teams while meeting US work eligibility requirements.

Qualifications

  • Candidates who are completing or have recently completed a PhD in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related discipline.
  • Strong grounding in theoretical and empirical research methods to address complex problems.
  • Experience with PyTorch or TensorFlow and familiarity with deep neural network architectures.
  • Experience with a range of ML models and algorithms, both neural and non-neural.

Responsibilities

  • Contribute cutting-edge research to ByteDance product evolution (e.g., Douyin, Capcut, and other apps).
  • Work on advanced science and technology in audio processing and generation (e.g., Dialogue Systems, Audio-Video Models, Speech Synthesis).
  • Research, model, design, develop and evaluate novel machine learning models and algorithms.
  • Collaborate with globally based researchers and engineering teams.

Skills

PyTorch
TensorFlow
Deep learning
Neural networks
C++
Python
Shell scripting

Education

PhD in Computer Science / Electrical Engineering / related

Tools

C/C++
Python
Shell scripting

Job description

ByteDance Seed is hiring researchers to push the frontiers of artificial general intelligence. The team explores MLLM, GenMedia, AI for Science, and Robotics, with global labs and opportunities in the United States, China, and Singapore.

Successful candidates will join a collaborative, cutting-edge environment. As a PhD holder, you will contribute to foundational models, speech and multimodal research, and work with world-class teams while meeting US work eligibility requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Graduate Research Scientist: Foundation Models & GenAI
Graduate Research Scientist: Foundation Models & GenAI

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Research Intern, Multimodal Speech & AI
Research Intern, Multimodal Speech & AI

ByteDance • San Jose (CA)

On-site
Health insurance
Life insurance
Wellbeing benefits
+3
Research Scientist, Multimodal Interaction & World Models
Research Scientist, Multimodal Interaction & World Models

ByteDance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
PhD Candidate: Multimodal AI Researcher
PhD Candidate: Multimodal AI Researcher

ByteDance • San Jose (CA)

On-site
USD 80,000 - 120,000
Graduate Research Scientist: Multimodal AI & Transformers
Graduate Research Scientist: Multimodal AI & Transformers

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+2
Speech AI Research Intern: Multimodal & Speech Tech
Speech AI Research Intern: Multimodal & Speech Tech

ByteDance • San Jose (CA)

On-site
Health insurance
Life insurance
Wellbeing benefits
+3
Research Scientist: AI Foundation Models & ML Systems
Research Scientist: AI Foundation Models & ML Systems

ByteDance • San Jose (CA)

On-site
USD 254,000 - 480,000
Research Scientist Graduate (Multimodal Interaction and World Model) - 2026 Start (PhD)
Research Scientist Graduate (Multimodal Interaction and World Model) - 2026 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+2
GenAI Research Scientist — Multimodal Vision & Audio
GenAI Research Scientist — Multimodal Vision & Audio

ByteDance • San Jose (CA)

On-site
USD 244,000 - 588,000
Medical, dental, and vision insurance
401(k) savings plan with company match
10 paid holidays and 10 sick days
+1
Multimodal Vision & AI Evaluation Researcher
Multimodal Vision & AI Evaluation Researcher

Bytedance • San Jose (CA)

Hybrid
USD 150,000 - 230,000