Staff Research Engineer - Multimodal Generative Modelling

SLAMcore

United Kingdom

Remote

GBP 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Synthesia, the world’s leading AI video platform, is seeking a Staff Research Engineer to shape the voice team’s roadmap and drive cross‑team research in multimodal, audio‑visual systems. You’ll own design, implementation and shipping of core components, collaborating with voice and video teams to deliver low‑latency, natural conversations and interactive experiences for enterprise customers.

Relentless curiosity, strong papers or open‑source contributions, and hands‑on experience with DL models

Qualifications

  • Novel ideas that advance multimodal interactive systems.
  • Strong understanding of generative modelling for sequential or multimodal data.
  • Hands‑on experience with Large Language Models or transformer architectures.
  • High proficiency in PyTorch, including distributed training and model optimisation.

Responsibilities

  • Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.
  • Propose novel multi-modal system architectures (especially text and voice).
  • Develop and evaluate streaming and conversational systems for low-latency interactive voice-video synthesis.
  • Design solutions that reinforce emotional expressiveness and natural interaction.
  • Implement and bring designs to life, from pretraining through post-training.
  • Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.
  • Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
  • Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
  • Curate new datasets to complement existing data.
  • Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality.
  • Ship models to production with optimised runtime to serve customers, and address their feedback thereafter.

Skills

Generative modelling
LLMs
PyTorch
Time-series modelling
Prototyping
Deep learning
Software engineering
Distributed training
Model optimisation

Tools

Distributed training
Model optimisation

Job description

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US.

As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations.

Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow.

About the role

Synthesia's long-term vision is to build the best human-interactive models: systems that don't just talk to people, but perceive and respond to them, reacting to a user's actions and emotions, not only their words. Today over 60,000 businesses rely on our platform, and the next leap in what we can offer them depends on models that combine text, audio, and video into a single real-time interactive experience.

As a Staff Research Engineer, you'll join the Voice team within our 40+ person R&D department, but your scope will extend well beyond voice. You'll help define and drive that broader vision across teams, proposing ambitious research directions and taking direct ownership of the design and implementation of its most critical components. You'll work directly with our voice lead and collaborate tightly with our video teams and other senior members of the org.

Concretely, you'll work on voice to voice models that produce text and voice simultaneously. Models that can reason, interrupt the user with back channeling and talk. Models that feel like you are having a natural conversation with, without the feeling of turn taking that is dominant in current speech to speech models. Your role would be to partner in defining a roadmap, implementing it and shipping the outcomes to product.

What you'll do
  • Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.

  • Propose novel multi-modal system architectures (especially text and voice).

  • Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.

  • Design solutions that reinforce emotional expressiveness and natural interaction.

  • Implement and bring designs to life, from pretraining through post-training.

  • Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.

  • Define new evaluation metrics for conversational systems, including latency‑aware and interaction‑based measurements.

  • Track the latest research in audio‑visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.

  • Curate new datasets to complement existing data.

  • Lead post‑training initiatives like DPO, fine‑tuning, and distillation to bring models to shipping quality.

  • Ship models to production with optimised runtime to serve customers, and address their feedback thereafter.

You’ll thrive in this role if you have
  • The ability to bring novel ideas and designs that advance the field of interactive multimodal systems.

  • Strong understanding of generative modelling, ideally applied to sequential or multimodal data.

  • Hands‑on experience with large language models or similar transformer‑based architectures.

  • High proficiency in PyTorch, including distributed training and model optimisation.

  • A solid grasp of time‑series modelling and tokenisation, preferably in the context of audio, speech, or video.

  • A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.

  • Proven experience training deep learning models end‑to‑end, from data preparation through evaluation.

  • Strong general software engineering skills, enabling contributions to a large, shared research infrastructure.

Particularly relevant experience
  • Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it.

  • Working on conversational or interactive systems where latency, responsiveness, and user experience were first‑class constraints, not afterthoughts.

  • Working on LLMs with large scale trainings leading to models with decent reasoning capabilities.

  • Owning a research problem end to end: from architecture proposal through pretraining, post‑training, and production deployment.

  • Collaborating across modalities or teams (e.g. audio and video, or research and product) to ship a unified system.

Bonus points for
  • Experience with real‑time or streaming architectures.

  • Familiarity with state‑of‑the‑art architectures in audio and speech generation, such as diffusion models, neural codecs, flow‑matching models, or autoregressive decoders.

  • Excellence in one or more of the following modalities: voice, text, video.

  • Evidence of original research contributions, such as publications or open‑source work at top‑tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Engineer: Multimodal Generative AI
Staff Research Engineer: Multimodal Generative AI

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000
Senior Applied Research Engineer - Video Team Synthesia Europe, Germany, Switzerland, UK
Senior Applied Research Engineer - Video Team Synthesia Europe, Germany, Switzerland, UK

Neura Market • United Kingdom

Remote
EUR 90,000 - 140,000
Competitive compensation (salary + 100
Fully remote from Europe or hybrid (UK
25 days annual leave + holidays
+2
Senior Applied Research Engineer - Video
Senior Applied Research Engineer - Video

Synthesia • Greater London

On-site
GBP 90,000 - 130,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • City Of London

On-site
GBP 85,000 - 120,000
Speech Research Scientist (TTS)
Speech Research Scientist (TTS)

ConnexAI • Manchester

On-site
GBP 70,000 - 95,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • Greater London

On-site
GBP 110,000 - 140,000
Senior Research Scientist | Multimodal Systems
Senior Research Scientist | Multimodal Systems

DeepL • Greater London

On-site
GBP 120,000 - 180,000
Senior Research Scientist | Multimodal Systems
Senior Research Scientist | Multimodal Systems

The Consensus • Greater London

On-site
GBP 110,000 - 170,000
Research Scientist - ASR
Research Scientist - ASR

ConnexAI • Manchester

On-site
GBP 70,000 - 120,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Speechmatics • Cambridge

On-site
GBP 90,000 - 140,000
Hybrid work model
Private Medical
Dental for you and family
+4