Staff Research Engineer, Multimodal Generative AI

synthesia

United States

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Synthesia is seeking a Staff Research Engineer to join the Voice team and drive a broader vision across teams. You will own design and implementation of core components for voice-to-voice models that synthesize text and speech in real time, collaborating with video teams and senior researchers.

Lead roadmap, prototype quickly, and ship end-to-end solutions from pretraining to production, while advancing multimodal interaction and latency-aware evaluation across products.

Qualifications

  • Strong understanding of generative modelling applied to sequential or multimodal data.
  • Hands-on experience with large language models or transformer-based architectures.
  • High proficiency in PyTorch, including distributed training and model optimisation.
  • Solid grasp of time-series modelling and tokenisation, preferably in audio, speech, or video.
  • Proven ability to prototype quickly, test hypotheses, and iterate efficiently.
  • Experience training deep learning models end-to-end, from data preparation through evaluation.
  • Strong general software engineering skills for a large research infrastructure.

Responsibilities

  • Shape roadmap to create new model capabilities and unlock functionality for customers.
  • Propose novel multi-modal architectures (text and voice).
  • Develop and evaluate streaming and conversational systems for low-latency synthesis.
  • Design solutions that reinforce emotional expressiveness and natural interaction.
  • Implement from pretraining through post-training and shipping to production.
  • Integrate and test architectures (neural codecs, diffusion, flow-matching) to enhance realism.
  • Define new evaluation metrics for conversational systems including latency-aware measures.
  • Track the latest research in audio-visual diffusion, autoregressive models, and multimodal LLMs.
  • Curate new datasets to complement existing data.
  • Lead post-training initiatives like DPO, fine-tuning, and distillation to shipping quality.
  • Ship models to production with optimised runtime and address feedback.

Skills

Generative modelling
Transformer architectures
PyTorch distributed training
Time-series modelling
Rapid prototyping
End-to-end training
Software engineering

Job description

Synthesia is seeking a Staff Research Engineer to join the Voice team and drive a broader vision across teams. You will own design and implementation of core components for voice-to-voice models that synthesize text and speech in real time, collaborating with video teams and senior researchers.

Lead roadmap, prototype quickly, and ship end-to-end solutions from pretraining to production, while advancing multimodal interaction and latency-aware evaluation across products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 100,000 - 130,000
Senior Multimodal AI Systems Engineer
Senior Multimodal AI Systems Engineer

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Full health and dental benefits
RRSP matching / 401(k)
+5
Senior Staff Research Scientist - Voice & Speech AI
Senior Staff Research Scientist - Voice & Speech AI

DeepL • London (KY)

On-site
USD 180,000 - 240,000
Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 140,000 - 230,000
Lead Real-Time Multimodal AI Scientist, Conversational/AGI
Lead Real-Time Multimodal AI Scientist, Conversational/AGI

Amazon • Bellevue (WA)

On-site
USD 180,000 - 260,000
Generative Audio Research Scientist — Vocal Synthesis (Equity)
Generative Audio Research Scientist — Vocal Synthesis (Equity)

Spotify AB • New York (NY)

On-site
USD 133,000 - 190,000
Health insurance
Six months parental leave
401(k) retirement plan
+3
Multimodal AI Research Scientist — Production-Focused
Multimodal AI Research Scientist — Production-Focused

descript • United States

On-site
USD 197,000 - 263,000
Research Engineer, Multimodal
Research Engineer, Multimodal

character • Redwood City (CA)

On-site
USD 100,000 - 150,000
Senior Multimodal AI Researcher (Audio) - Flexible Work
Senior Multimodal AI Researcher (Audio) - Flexible Work

Dolby • Atlanta (GA)

On-site
USD 141,000 - 170,000
Flex Work