Staff Research Engineer: Multimodal Generative AI

SLAMcore

United Kingdom

Remote

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Synthesia, the world’s leading AI video platform, is seeking a Staff Research Engineer to shape the voice team’s roadmap and drive cross‑team research in multimodal, audio‑visual systems. You’ll own design, implementation and shipping of core components, collaborating with voice and video teams to deliver low‑latency, natural conversations and interactive experiences for enterprise customers.

Relentless curiosity, strong papers or open‑source contributions, and hands‑on experience with DL models

Qualifications

  • Novel ideas that advance multimodal interactive systems.
  • Strong understanding of generative modelling for sequential or multimodal data.
  • Hands‑on experience with Large Language Models or transformer architectures.
  • High proficiency in PyTorch, including distributed training and model optimisation.

Responsibilities

  • Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.
  • Propose novel multi-modal system architectures (especially text and voice).
  • Develop and evaluate streaming and conversational systems for low-latency interactive voice-video synthesis.
  • Design solutions that reinforce emotional expressiveness and natural interaction.
  • Implement and bring designs to life, from pretraining through post-training.
  • Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.
  • Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
  • Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
  • Curate new datasets to complement existing data.
  • Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality.
  • Ship models to production with optimised runtime to serve customers, and address their feedback thereafter.

Skills

Generative modelling
LLMs
PyTorch
Time-series modelling
Prototyping
Deep learning
Software engineering
Distributed training
Model optimisation

Tools

Distributed training
Model optimisation

Job description

Synthesia, the world’s leading AI video platform, is seeking a Staff Research Engineer to shape the voice team’s roadmap and drive cross‑team research in multimodal, audio‑visual systems. You’ll own design, implementation and shipping of core components, collaborating with voice and video teams to deliver low‑latency, natural conversations and interactive experiences for enterprise customers.

Relentless curiosity, strong papers or open‑source contributions, and hands‑on experience with DL models

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000
Senior Research Engineer — Real-Time Voice Synthesis
Senior Research Engineer — Real-Time Voice Synthesis

Synthesia • United Kingdom

Remote
GBP 75,000 - 110,000
Senior Staff Research Scientist: Real-Time Voice & Multilingual AI
Senior Staff Research Scientist: Real-Time Voice & Multilingual AI

DeepL GmbH • Greater London

Hybrid
GBP 120,000 - 180,000
Voice AI Researcher: Expressive TTS & Audio Synthesis
Voice AI Researcher: Expressive TTS & Audio Synthesis

DNEG • Greater London

On-site
Senior Staff Research Scientist - Real-Time Voice AI
Senior Staff Research Scientist - Real-Time Voice AI

DeepL • Greater London

Hybrid
GBP 110,000 - 160,000
Senior Video Diffusion Research Engineer
Senior Video Diffusion Research Engineer

Synthesia • Greater London

On-site
GBP 90,000 - 130,000
Speech Research Scientist (TTS)
Speech Research Scientist (TTS)

ConnexAI • Manchester

On-site
GBP 70,000 - 95,000
GenAI Applied Scientist – Multimodal Speech & Language
GenAI Applied Scientist – Multimodal Speech & Language

Amazon Science • Cambridge

On-site
GBP 90,000 - 150,000
Research Scientist: Real-time Multimodal Generative AI
Research Scientist: Real-time Multimodal Generative AI

Google DeepMind • Greater London

On-site
GBP 90,000 - 150,000
Senior Staff Research Scientist, Real-Time Voice AI
Senior Staff Research Scientist, Real-Time Voice AI

AI Startups UK • Greater London

Hybrid
GBP 103,000 - 148,000