ML Engineer, TTS & Voice AI

Cantina Labs

Greater London

On-site

GBP 148,600 - 163,460

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
PTO 42 days
Parental leave
401(k)
Free lunch

Job summary

Cantina is seeking a Research / ML Engineer for our Speech Team to build state-of-the-art speech systems end-to-end, from data specs to production inference. You’ll drive the model data eval flywheel for TTS and related tasks, partnering with research, data, and infra to ship reliable, cost-aware models.

You’ll lead small research projects, design experiments, and develop tooling, contributing to safety and responsible AI while pushing the boundaries of voice cloning and controllable TTS.

Qualifications

  • Experience with large-scale audio models and data.
  • Strong understanding of transformer/diffusion models and audio language modelling.
  • Expertise in multi-node distributed training and PyTorch.
  • Proven ability to ship production-ready speech models.

Responsibilities

  • Architect, implement and fine-tune large-scale speech models.
  • Lead small research projects and collaborate on larger goals.
  • Design and run experiments to advance model quality.
  • Develop dev tooling to boost team productivity.
  • Contribute across the stack from low-level optimizations to model design.
  • Define data needs and oversee data acquisition, curation and labeling quality.
  • Design automated evaluations and monitor model safety and bias.
  • Harden training to inference pipelines and manage latency and cost.
  • Collaborate with infra to scale training on cloud fleets and productionize models.
  • Contribute to safety guardrails and misuse mitigation.

Skills

Large-scale audio models
Transformer architectures
Multi-node distributed training
PyTorch
Production-quality code
Voice cloning / speech control
Publications / open source

Tools

CUDA
Triton
Distributed training tooling
Python

Job description

Cantina is seeking a Research / ML Engineer for our Speech Team to build state-of-the-art speech systems end-to-end, from data specs to production inference. You’ll drive the model data eval flywheel for TTS and related tasks, partnering with research, data, and infra to ship reliable, cost-aware models.

You’ll lead small research projects, design experiments, and develop tooling, contributing to safety and responsible AI while pushing the boundaries of voice cloning and controllable TTS.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, TTS
Machine Learning Engineer, TTS

Cantina Labs • Greater London

On-site
GBP 148,000 - 164,000
Equity
Health insurance
PTO 42 days
+3
Founding ML Engineer: Audio AI & Model Evaluation
Founding ML Engineer: Audio AI & Model Evaluation

Cantina Labs • Greater London

On-site
GBP 150,000 - 166,000
Competitive salary and generous equity
Medical, dental, and vision insurance
42 days of paid time off including 15.
+2
Senior Staff Research Scientist: Real-Time Voice & Multilingual AI
Senior Staff Research Scientist: Real-Time Voice & Multilingual AI

DeepL GmbH • Greater London

Hybrid
GBP 120,000 - 180,000
Staff Research Engineer: Multimodal Generative AI
Staff Research Engineer: Multimodal Generative AI

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000
Research Engineer - Llms & Generative Audio
Research Engineer - Llms & Generative Audio

Harnham • Greater London

Hybrid
GBP 90,000 - 150,000
Voice AI Researcher: Expressive TTS & Audio Synthesis
Voice AI Researcher: Expressive TTS & Audio Synthesis

DNEG • Greater London

On-site
Senior LLM Engineer - Multimodal Audio & Voice AI
Senior LLM Engineer - Multimodal Audio & Voice AI

Harnham • Slough

Hybrid
GBP 108,000 - 132,000
Hybrid work model
London-based office
Machine Learning - (Speech) - Contract
Machine Learning - (Speech) - Contract

microTECH Global LTD • Egham

On-site
GBP 113,000 - 152,000
Senior Engineering Manager, Real-Time AI Voice
Senior Engineering Manager, Real-Time AI Voice

AI Startups UK • Greater London

Hybrid
GBP 110,000 - 160,000
Hybrid work schedule
Virtual Shares
Team events and Hack Fridays
+2
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000