Senior Audio AI Engineer – TTS / Speech Synthesis

Awarri

United Kingdom

Hybrid

GBP 60,000 - 80,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Be part of a pioneering initiative
Work on impactful projects
Collaborate with a passionate team

Job summary

A leading AI technology firm in the United Kingdom is seeking a Senior Audio AI Engineer to advance Text-to-Speech systems focused on African languages. This role requires deep experience in speech technologies and aims to improve the quality of audio AI in cultural contexts. You'll optimize neural models, evaluate TTS systems, and ensure production readiness. Ideal candidates have a strong proficiency in Python and experience with machine learning frameworks. Join a mission-driven team making a significant impact.

Qualifications

  • 3+ years of experience developing and deploying TTS or speech generation systems.
  • Deep knowledge of at least one neural TTS architecture and related vocoders.
  • Experience with audio processing tools like librosa or Praat.

Responsibilities

  • Optimize neural TTS models for prosody, pacing, and expressiveness.
  • Evaluate and fine-tune neural vocoders to match desired audio quality.
  • Lead evaluations of TTS quality using objective and subjective measures.

Skills

Proficiency in Python
Experience with TypeScript
Background in machine learning frameworks
Knowledge of database management using MongoDB
Experience in scalable AI-driven applications

Tools

TensorFlow
PyTorch
FastAPI
Flask
MySQL

Job description

At Awarri, our mission is to enable the development and adoption of frontier technology across Africa, starting in Nigeria. We are building inclusive AI technologies—from LLMs to speech models—that reflect and empower African languages and cultural contexts.

Why Join Awarri?

  • Be part of a pioneering initiative shaping the future of AI in Africa.
  • Work on impactful projects that center real-world representation and inclusivity.
  • Collaborate with a passionate, globally distributed team of engineers, linguists, and researchers.

As a Senior Audio AI Engineer at Awarri, you will play a pivotal role in advancing the naturalness and quality of our Text-to-Speech (TTS) systems, focused on African languages and accents. We’re seeking an engineer who understands the intricacies of prosody, rhythm, and speech alignment—and is excited to push the boundaries of audio AI in a meaningful cultural context.

This role is best suited for a specialist with deep experience in speech technologies and a passion for building expressive, production-ready TTS models. You’ll be joining a collaborative, mission-driven team dedicated to shaping the future of generative audio systems in Africa.

Responsibilities
Model Development & Fine-Tuning
  • Optimize neural TTS models for prosody, pacing, and expressiveness (e.g., Tacotron 2, FastSpeech 2, Glow-TTS, VITS).
  • Improve duration prediction and phoneme-to-frame alignment using forced aligners or prosody-aware training.
  • Incorporate punctuation and linguistic markers into the model pipeline to improve natural flow.
  • Implement and fine-tune transformer-based architectures for speech synthesis and text-to-speech tasks.
Audio Engineering & Vocoder Optimization
  • Evaluate and fine-tune neural vocoders (e.g., HiFi-GAN, WaveGlow) to match desired voice characteristics and audio quality.
  • Identify and correct audio artifacts or inconsistencies in generated speech.
  • Optimize speech processing pipelines for efficiency and real-time performance.
Evaluation & Iteration
  • Lead both objective (e.g., duration errors, pitch contours) and subjective (e.g., MOS scoring) evaluations of TTS quality.
  • Collaborate with linguistic teams to benchmark pronunciation accuracy in Nigerian languages.
  • Develop automated testing frameworks to validate speech synthesis quality at scale.
Deployment & Production Readiness
  • Prepare the TTS system for product integration by improving inference speed and robustness.
  • Support the deployment of models across various platforms (cloud, mobile, embedded).
  • Optimize model inference using VLLM for efficient deployment.
  • Build APIs and backend services for TTS deployment using FastAPI and Flask.
  • Implement and manage data pipelines and storage solutions using MongoDB and MySQL.
Technical Skills & Requirements
  • Proficiency in Python and TypeScript for model development and backend integration.
  • Experience with transformer-based models for speech synthesis and NLP.
  • Strong background in machine learning frameworks such as TensorFlow or PyTorch.
  • Experience in designing scalable AI-driven applications.
  • Familiarity with FastAPI, Flask, and cloud-based deployment environments.
  • Knowledge of database management using MongoDB and MySQL.
Your Experience
Technical Expertise
  • 3+ years of experience developing and deploying TTS or speech generation systems (bonus for low-resource languages).
  • Deep knowledge of at least one neural TTS architecture and related vocoders.
  • Proficiency with PyTorch, TensorFlow, or JAX for building and training models.
  • Experience with audio processing tools (e.g., librosa, Praat, torchaudio).
Linguistic & Cultural Sensitivity
  • Experience working with multilingual or low-resource speech data.
  • Familiarity with phonetics/phonology, especially as it relates to prosody and rhythm.
Engineering Workflow
  • Experience building scalable training and evaluation pipelines.
  • Ability to debug complex model behavior and iterate quickly toward product quality.
  • Comfort working remotely and asynchronously with interdisciplinary teams.
Nice to Have
  • Prior work on African language speech systems or expressive TTS in non-English languages.
  • Interest in linguistic or cultural technology in the African context.
  • Contributions to open-source TTS or audio AI tools.
  • Experience with emotion modeling or speaker adaptation.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Audio AI Engineer: Remote TTS & Speech Synthesis
Senior Audio AI Engineer: Remote TTS & Speech Synthesis

Awarri • United Kingdom

Hybrid
GBP 60,000 - 80,000
Be part of a pioneering initiative
Work on impactful projects
Collaborate with a passionate team
Machine Learning Engineer
Machine Learning Engineer

ConnexAI • Manchester

On-site
GBP 70,000 - 110,000
Speech Research Scientist (TTS)
Speech Research Scientist (TTS)

ConnexAI • Manchester

On-site
GBP 65,000 - 90,000
Voice AI Researcher: Expressive TTS & Audio Synthesis
Voice AI Researcher: Expressive TTS & Audio Synthesis

DNEG • Greater London

On-site
Confidential
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000
Research Engineer - LLMs & Generative Audio
Research Engineer - LLMs & Generative Audio

Harnham • Slough

On-site
GBP 108,000 - 132,000
Hybrid work model
London-based office
Account Manager - Speech & AI Technologies
Account Manager - Speech & AI Technologies

PC Games Insider • United Kingdom

Hybrid
GBP 60,000 - 90,000
Private healthcare options
Life assurance
Income protection
+3
Software Engineer, Data Infrastructure & Acquisition - Belfast, United Kingdom
Software Engineer, Data Infrastructure & Acquisition - Belfast, United Kingdom

Speechify • Belfast City District

On-site
GBP 89,638 - 112,048
Competitive salary
Asynchronous culture
Life-changing product
+1
Machine Learning - (Speech) - Contract
Machine Learning - (Speech) - Contract

microTECH Global LTD • Egham

On-site
GBP 113,991 - 151,988
Research Scientist - ASR
Research Scientist - ASR

ConnexAI • Manchester

On-site
GBP 70,000 - 90,000