Research Scientist - Video Diffusion

Nuance Labs

Seattle (WA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An innovative AI startup in Seattle is looking for experts to develop a groundbreaking human foundation model that integrates text, speech, and emotional signals in real time. Candidates should have a PhD or equivalent experience with a strong background in training audio generation models and deep learning. You will work alongside a top-tier research team to create lifelike avatars that understand and respond to human nuances, helping bridge the emotional gap in AI. This role offers a collaborative and fast-paced environment where your contributions will directly impact the development of cutting-edge technology.

Qualifications

  • PhD (or equivalent) with experience in training audio generation models.
  • Strong understanding of deep learning and ML pipeline.
  • Ability to solve blank-page problems independently.
  • Collaborative mindset and ability to work with researchers across domains.

Responsibilities

  • Develop the first human foundation model integrating text, speech, and body language.
  • Create lifelike avatars capable of nuanced responses.
  • Innovate in real-time multimodal interaction technology.
  • Translate research breakthroughs into practical, real-world AI products.
  • Contribute to a fast-paced, in-person Seattle team.

Skills

Training speech synthesis models
Deep learning
ML pipeline management
Clean code writing
Collaboration with diverse teams
Benchmarking
Evaluation

Education

PhD or equivalent experience in relevant fields

Job description

Nuance Labs is an early-stage deep tech startup. We’re building the first real-time human foundation model — unifying text, speech, and vision — to make AI socially and emotionally intelligent. Imagine an AI that can understand a quirked eyebrow, a shift in tone, or a hesitant pause, and respond in a way that feels truly human.

Key Facts

$10M seed round backed by Accel, South Park Commons, Lightspeed, and top angels including Synthesia’s former CPO.

A world-class team of PhDs from MIT, UW, and Oxford with decades of industry experience at Apple and Meta, advancing real-time avatars from cutting-edge research to products used by millions.

In-person collaboration, 5 days a week at Seattle HQ

This is for you, if

Have a PhD (or equivalent experience) in training speech synthesis models (text-to-speech, speech-to-speech, etc.), training audio generation models, or related fields, with a track record of pushing the research frontier

Know deep learning inside out and can run the whole ML pipeline, from data wrangling and rapid prototyping to large-scale training, benchmarking, and evaluation

Love blank-page problems, chart your own course, and make progress without waiting for someone to hand you a task list

Move quickly from research breakthroughs to practical, real-world applications

Write code that’s clean enough your future self will thank you for

Play well with other brilliant minds from different domains

What you’ll be building

The first human foundation model that operates across text, speech, facial expression, and body language in real time. This unified model:

Understands fine-grained human signals — from a quirked eyebrow to a subtle change in voice — and infers meaning in context

Generates lifelike, responsive avatars whose expressions, gestures, and tone evolve frame-by-frame to deliver genuine responses

The landscape is ripe for innovation.While voice AI systems have made great strides in capturing prosody, and avatar platforms can generate compelling visuals, existing solutions remain fragmented. Real-time, multimodal interaction — where voice, facial expression, and contextual perception converge — is still an unsolved problem. This role offers the rare opportunity to shape foundational technology in a space where the boundaries are still being defined.

Why this team

We’re research scientists who’ve spent years advancing AI avatar and audio-visual generation — publishing at top conferences and shipping ultra-low-latency ML products to millions. We combine frontier research with the ruthless engineering needed for consumer-grade, real-time systems.

To apply, email us with your CV and a short note on why your background is a great fit for this role.

Send application or questions to careers@nuancelabs.ai

VALUES
  • Radical transparency: We communicate openly so everyone can make informed decisions.
  • Relentless speed: We bias towards action, iterate fast, and learn quickly.
  • Doing right by people: Integrity and respect are not negotiable.
  • Being together fuels our energy and accelerates our problem-solving.
Join us in bridging the emotional gap of artificial intelligence
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - Data
ML Engineer - Data

Nuance Labs • Seattle (WA)

On-site
USD 90,000 - 120,000
Member of Technical Staff — ML Infra (Data)
Member of Technical Staff — ML Infra (Data)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account contributions
15 days PTO plus holidays
Lunch, drinks, and snacks provided
Research Assistant
Research Assistant

Nuance Labs • Seattle (WA)

Hybrid
USD 28,000 - 41,000
Germany Senior / Lead Research Scientist - Germany
Germany Senior / Lead Research Scientist - Germany

Inworld AI • Germany (OH)

On-site
USD 136,000 - 205,000
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Conversational Modelling Research Engineer
Conversational Modelling Research Engineer

Tavus • San Francisco (CA)

Hybrid
USD 120,000 - 230,000
Member of Technical Staff — Pretraining Infra (Experienced)
Member of Technical Staff — Pretraining Infra (Experienced)

Nuance Labs • Seattle (WA)

On-site
USD 300,000 - 400,000
Health contributions
15 days of PTO
Commuter benefits
+1
Member of Technical Staff — RL Research (New PhD Grad)
Member of Technical Staff — RL Research (New PhD Grad)

Nuance Labs • Seattle (WA)

On-site
USD 250,000 - 350,000
Health savings account contributions
15 days of PTO
Lunch, drinks, and snacks provided
Member of Technical Staff — RL Research (Experienced)
Member of Technical Staff — RL Research (Experienced)

Nuance Labs • Seattle (WA)

On-site
USD 300,000 - 500,000
Health Savings Account plan
15 days PTO plus holidays
Lunch and snacks provided
Remote Voice AI Model Research Engineer
Remote Voice AI Model Research Engineer

Deepslate • Germany (OH)

Remote
USD 90,000 - 150,000
True R&D Freedom
Competitive Compensation
Serious Compute
+1