Machine Learning Engineer, TTS

Cantina Labs

Greater London

On-site

GBP 148,600 - 163,460

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
PTO 42 days
Parental leave
401(k)
Free lunch

Job summary

Cantina is seeking a Research / ML Engineer for our Speech Team to build state-of-the-art speech systems end-to-end, from data specs to production inference. You’ll drive the model data eval flywheel for TTS and related tasks, partnering with research, data, and infra to ship reliable, cost-aware models.

You’ll lead small research projects, design experiments, and develop tooling, contributing to safety and responsible AI while pushing the boundaries of voice cloning and controllable TTS.

Qualifications

  • Experience with large-scale audio models and data.
  • Strong understanding of transformer/diffusion models and audio language modelling.
  • Expertise in multi-node distributed training and PyTorch.
  • Proven ability to ship production-ready speech models.

Responsibilities

  • Architect, implement and fine-tune large-scale speech models.
  • Lead small research projects and collaborate on larger goals.
  • Design and run experiments to advance model quality.
  • Develop dev tooling to boost team productivity.
  • Contribute across the stack from low-level optimizations to model design.
  • Define data needs and oversee data acquisition, curation and labeling quality.
  • Design automated evaluations and monitor model safety and bias.
  • Harden training to inference pipelines and manage latency and cost.
  • Collaborate with infra to scale training on cloud fleets and productionize models.
  • Contribute to safety guardrails and misuse mitigation.

Skills

Large-scale audio models
Transformer architectures
Multi-node distributed training
PyTorch
Production-quality code
Voice cloning / speech control
Publications / open source

Tools

CUDA
Triton
Distributed training tooling
Python

Job description

About Cantina

Cantina is a new social platform founded by Sean Parker with the most advanced AI character creator. Our bots are lifelike, social creatures that can interact wherever people are online—across voice, video, and text. Create yourself, imagine someone new, or choose from thousands of characters to share infinitely scalable, personalized content and seamless group chat.

If you’re excited about how AI can shape creativity and social interaction, come help us build what’s next.

About The Role

We’re looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech systems end-to-end—from data specs through production inference. You’ll drive the model data eval flywheel for TTS and adjacent tasks (voice cloning, controllable TTS, voice conversion and more), partnering closely with research, data, and infra to ship fast, reliable, and cost-aware models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

What You’ll Do
  • Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) for large-scale speech models.
  • Project Leadership: Independently lead small research projects while collaborating on larger team initiatives.
  • Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.
  • Tool Development: Develop and improve dev tooling to enhance team productivity.
  • Full-Stack Contribution: Contribute to the entire stack, from low-level optimizations to high-level model design.
  • Data Ownership: Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
  • Rigorous Evaluation: Design automated objective/subjective evaluations—listening tests, SV/WER/ASR-based metrics, robustness & bias checks, and red-team studies.
  • Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
  • GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.
  • Safety & Responsibility: Contribute to safety/consent guardrails and to misuse/abuse mitigation for responsible speech technology.
What You’ll Bring
  • Exceptional research/development experience with large scale audio models (>3B models and >500k hours data).
  • Exceptional understanding and hands‑on experience with transformer architectures and/or diffusion models (inc. distillation and streaming) and/or audio language modelling.
  • Strong experience with multi-node and multi-gpu distributed model training.
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production quality code.
  • Shipped large scale speech/audio models to production.
  • Background in working with large-scale ML data.
  • Ability to iterate on data,, and triangulate quality using subjective and objective signals.
  • Notable publications and/or open source contributions in speech/audio/ML.
  • Experience with voice-cloning, speech-control, voice-generation.
Preferred Experience
  • Shipped large scale speech/audio models (TTS/VC/ASR) to production.
  • Work on large-scale ML systems.
  • Experience with audio language modelling, transformer architectures.
  • Experience with voice-cloning, speech-control, voice-generation.
  • Background in processing large-scale ML data.
  • Publications or notable open-source in speech/audio/ML.
Compensation

The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits For U.S.-based Roles
  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Data & ML Infrastructure for Video Models
Member of Technical Staff, Data & ML Infrastructure for Video Models

Cantina Labs • Greater London

On-site
GBP 150,000 - 196,000
Medical insurance
Paid time off
Parental leave
+4
ML Engineer, TTS & Voice AI
ML Engineer, TTS & Voice AI

Cantina Labs • Greater London

On-site
GBP 148,000 - 164,000
Equity
Health insurance
PTO 42 days
+3
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Speechmatics • Cambridge

On-site
GBP 90,000 - 140,000
Hybrid work model
Private Medical
Dental for you and family
+4
ML Data & Platform Engineer
ML Data & Platform Engineer

Speechmatics • Cambridge

On-site
GBP 70,000 - 110,000
Private Medical
Dental for you and your family
Global working opportunities
+5
ML Data & Platform Engineer
ML Data & Platform Engineer

Speechmatics • Greater London

On-site
GBP 90,000 - 120,000
Flexible hours
Company lunches
Birthday celebrations
+7
ML Data & Platform Engineer
ML Data & Platform Engineer

Speechmatics • United Kingdom

On-site
GBP 90,000 - 120,000
Private Medical
Dental for family
Global opportunities
+4
Software Engineer, Platform - Newcastle, United Kingdom Newcastle, United Kingdom
Software Engineer, Platform - Newcastle, United Kingdom Newcastle, United Kingdom

Speechify • Newcastle upon Tyne

Hybrid
GBP 85,000 - 120,000
Dynamic environment
Autonomy
Impact on AI/audio
+2
Software Engineer, Platform - Manchester, United Kingdom Manchester, United Kingdom
Software Engineer, Platform - Manchester, United Kingdom Manchester, United Kingdom

Speechify • Manchester

Hybrid
GBP 90,000 - 120,000
Fully distributed team
Competitive compensation
Impactful product
Software Engineer, Platform - Nottingham, United Kingdom Nottingham, United Kingdom
Software Engineer, Platform - Nottingham, United Kingdom Nottingham, United Kingdom

Speechify • Nottingham

Hybrid
GBP 70,000 - 90,000
Dynamic environment
Autonomy to own projects
High-impact product
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000