Conversational Modelling Research Engineer

Tavus

United States

Hybrid

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tavus is seeking an AI Researcher to join their core AI team in the United States. The role involves conducting advanced research on Large Multimodal Models, particularly for Conversational Avatars. Candidates should have a PhD (or near completion) and hands-on experience with generative models and deep learning. The position offers a preferred location in San Francisco or London, with potential remote work for exceptional candidates. Join us to push the boundaries of AI technology.

Qualifications

  • Hands-on experience with Large Multimodal Models.
  • Experience in fine-tuning/adapting VLMs for control and conditioning.
  • Solid background in deep learning and foundation models.

Responsibilities

  • Conduct research on Large Multimodal Models for Conversational Avatars.
  • Develop methods to model verbal and non-verbal conversation aspects.
  • Partner with Applied ML team to transition research from prototype to production.

Skills

Large Multimodal Models
Generative language models
Deep learning
PyTorch

Education

PhD or equivalent research experience

Job description

The Role

We’re looking for an AI Researcher to join our core AI team and push the boundaries of Foundation Multimodal Conversational Models. If you thrive in fast-moving startup environments, enjoy experimenting with new ideas, and love seeing your work come to life in production then you’ll feel right at home.

Your Mission
  • Conduct research on Large Multimodal Models in the context of Conversational Avatars (e.g. Neural Avatars, Talking-Heads).
  • Develop methods to model both verbal and non-verbal aspects of conversation, adapting and controlling avatar behavior in real time, with low-latency.
  • Experiment with fine-tuning, adaptation, and conditioning techniques to make AudioVisual Multimodal Models, more expressive, controllable, and task-specific.
  • Partner with the Applied ML team to take research from prototype to production.
  • Stay up to date with cutting‑edge advancements — and help define what comes next.
You’ll Be Great At This If You Have:
  • A PhD (or near completion) in a relevant field, or equivalent research experience.
  • Hands‑on experience with Large Multimodal Models and a strong foundation in generative (language) models. This could be in the context of tasks such as VQA, Audio/Video understanding tasks, captioning behavioral analysis, Translation tasks, Speech to Speech systems.
  • Experience in fine‑tuning/adapting VLMs for control, conditioning, or downstream tasks.
  • Solid background in deep learning and foundation modes.
  • Strong PyTorch skills and comfort building deep learning pipelines.
Nice‑to‑Haves
  • Knowledge of large‑scale model training and optimization.
  • Experience in duplex‑conversational model.
  • Broader understanding of generative AI across modalities.
  • Exposure to software development best practices.
  • A flexible, experimental mindset i.e. comfortable working across research and engineering.
  • (Bonus) Publications at EMNLP, COLING, NeurIPS, ICLR, CVPR, ICCV.
Location

Preferred: San Francisco (hybrid) or London (office opening soon).

Remote within the U.S. or Europe available for exceptional candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Conversational Modelling Research Engineer
Conversational Modelling Research Engineer

Tavus • San Francisco (CA)

Hybrid
USD 120,000 - 230,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Multimodal Conversational Research Engineer
Multimodal Conversational Research Engineer

Tavus • United States

Hybrid
USD 100,000 - 150,000
Research Engineer, Multimodal
Research Engineer, Multimodal

character • Redwood City (CA)

On-site
USD 100,000 - 150,000
Research Scientist - Audio [33363]
Research Scientist - Audio [33363]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Audio [33341]
Research Scientist - Audio [33341]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Multimodal AI Researcher
Multimodal AI Researcher

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Research Engineer
Research Engineer

Nace.AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Research Scientist - Video Diffusion
Research Scientist - Video Diffusion

Nuance Labs • Seattle (WA)

On-site
USD 150,000 - 210,000
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3