Doctoral Researcher in Speech and Language Technology

Aalto

Espoo

Hybrid

EUR 30,000 - 40,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Aalto University in Espoo, Finland, invites applications for a Doctoral Researcher in Speech and Language Technology. The project focuses on detecting AI-generated audio, with work across watermarking, detection robustness, and attribution. You will develop detectors, build datasets, and publish results in top venues.

Under supervision by Dr. Lauri Juvela, you will collaborate internationally, use HPC resources, and contribute to open-source software and models.

Qualifications

  • Master’s degree or near completion in related field (speech, audio processing, ML).
  • Strong interest in doctoral research on audio deepfake detection, watermarking, and generative AI.
  • Good programming skills, especially Python and ML tools.
  • Fluent in English; Finnish language not required.

Responsibilities

  • Detect generated audio content from multiple sources.
  • Develop datasets of real and generated speech and audio.
  • Adapt detector baselines for multi-objective detection.
  • Publish research and contribute to open-source releases.

Skills

Audio deepfake detection
Watermarking
Source tracing
Generative AI
Python
ML tools
Experimental evaluation

Education

Master's degree in related field
Near completion of Master's degree

Tools

Python
Deep learning frameworks (TensorFlow/PyTorch)

Job description

## Doctoral Researcher in Speech and Language TechnologyApply: Otaniemi, Espoo, Finland: Full time: Posted 2 Days Ago: R48125*Aalto University is where science and art meet technology and business. We shape a sustainable future by making research breakthroughs in and across our disciplines, sparking the game changers of tomorrow and creating novel solutions to major global challenges. Our community is made up of 16 000 students and 5 200 employees, including 446 professors. Our campus is in Espoo, Greater Helsinki, Finland. Diversity is part of who we are, and we actively work to ensure our community’s diversity and inclusiveness. This is why we warmly encourage qualified candidates from all backgrounds to join our community.*The research will be carried out at the **Department of Information and Communications Engineering, DICE**, at Aalto University, Finland. The project environment offers excellent infrastructure for deep learning, speech, and audio research, including Aalto University’s large-scale scientific computing cluster with CPU and GPU nodes, access to CSC’s national computing infrastructure including LUMI, and the Aalto Acoustics Lab with anechoic chambers, listening rooms, and audio measurement equipment.**We are now looking for a Doctoral Researcher in Speech and Language Technology**Are you excited about speech and audio AI, deepfake detection, and the question of how we can identify the origin of AI-generated content? We are looking for a **Doctoral Researcher** to join the funded project **COWAMA: Content-based watermarking for AI-generated audio**, led by Assistant Professor Lauri Juvela.The rapid development of generative AI has made it possible to create realistic speech and music, but it has also created risks related to misinformation, malicious deepfakes, copyright, ownership, attribution, and accountability. The project addresses these challenges by developing methods for detecting the origin and authenticity of generated and watermarked audio data, improving content-based audio watermarking, and evaluating robustness against removal, spoofing, adversarial attacks, and realistic downstream processing.**Your role and goals**As the Doctoral Researcher, you will take primary responsibility for **detecting generated audio content from multiple sources**, and work jointly with the Postdoctoral Researcher on **robustness, evaluation, and generalization**.Your work will include:* Developing methods for detecting the authenticity and origin of speech and audio content generated by multiple models and containing different watermarking methods.* Creating datasets of diverse real and generated speech and audio, including content from different generative models and multiple watermarking methods.* Adapting and improving detector baselines for multi-objective detection involving deepfake detection, content origin detection, and watermark identification.* Developing explainable and localized detection methods for watermark and deepfake attribution, especially when generated content is mixed with other audio sources.* Studying robustness against signal-processing attacks, generative resynthesis attacks, adversarial attacks, and realistic downstream processing such as mixing.* Publishing research in leading journals and conferences in speech, audio, and machine learning, and contributing to open-source releases of software, trained models, and reproducible research outputs.The research methods will include experiment design, software implementation, running deep learning experiments on high-performance computing infrastructure, and evaluation using objective metrics and subjective listening tests.**Your network and team**You will be supervised by **Assistant Professor Lauri Juvela**, who leads the Speech Synthesis research group at Aalto University. The group works on deep generative models for speech and audio, speech synthesis, differentiable signal processing, deepfake detection, and watermarking.In addition to the Aalto speech groups and the Postdoctoral Researcher in the project, you will work with an international collaboration network including Prof. Junichi Yamagishi at the National Institute of Informatics in Japan, Prof. Xavier Serra at Universitat Pompeu Fabra in Spain, Prof. Gustav Eje Henter at KTH Royal Institute of Technology in Sweden.The collaboration network has a strong background in speech and audio synthesis, deep generative models, and deepfake detection, positioning the project to make an impact in watermarking and source tracing for AI-generated audio.**Your experience and ambitions**We are looking for a curious and motivated early-career researcher who wants to develop expertise in speech, audio, machine learning, and trustworthy generative AI. To succeed in this role, you should have:* A master’s degree, or be close to completing a master’s degree, in speech or audio processing, machine learning, signal processing, computer science, electrical engineering, or a related field.* Strong interest in doctoral research on audio deepfake detection, watermarking, source tracing, and generative AI.* Good programming skills, preferably including Python and modern machine learning tools.* Basic knowledge of deep learning, signal processing, speech/audio processing, or statistical machine learning.* Interest in working with speech and audio datasets, detector models, generative audio models, and experimental evaluation.* Motivation to publish scientific articles and contribute to open-source and reproducible research. [* Ability to work both independently and collaboratively in an international research environment.* **Fluency in English is required. Finnish language is not required.**If you are chosen for this position, you will apply for the study right in doctoral studies at Aalto University School of Electrical Engineering. Thus, please see the student information and admission criteria at https://www.aalto.fi/en/study-options/aalto-doctoral-programme-in-electrical-engineering.**What we offer*** A doctoral research position in a timely and socially meaningful field: improving transparency, traceability, and accountability for AI-generated speech and audio.* The opportunity to work on technical methods that support safer use of generative AI, including deepfake detection, watermark attribution, and robustness evaluation.* Excellent computing and audio research infrastructure, including CPU/GPU clusters, access to CSC and LUMI, FIN-CLARIN resources, and the Aalto Acoustics Lab.* A supportive research team and supervision by Assistant Professor Lauri Juvela.* International collaboration opportunities with leading researchers in speech synthesis, music technology, deepfake detection, and generative audio.* A strong open-science environment: the project aims to publish in leading venues and release software source code and trained models to support reproducibility and FAIR data management.* Great possibilities for competence development and learning, including professional development opportunities, staff training, and development projects based on your interests and needs.* A culture guided by responsibility, courage, and collaboration, where equality and inclusion support curiosity, innovation, collaboration, and wellbeing.Our vast array of professional development opportunities means you will grow and learn, having the chance to participate actively in staff training and development projects based on your interests and needs. The starting salary for this position is 3143 €/month and increases after a mid-term evaluation. The position is fixed term and follows school’s standard 2+2 model. It will be made initially for two years, with a six-month probationary period, and extended by two further years after a successful mid-term review, giving a total duration of four years. The position starts in January 2027 or as mutually agreed.We value work-life balance and well-being in all aspects of life. We work in a hybrid model, with the primary workplace located at the Otaniemi Campus in Espoo, Finland. Life on the revitalized campus is vibrant, featuring stunning architecture, tranquil nature, and a variety of cafes, restaurants, and services, all complemented by excellent public transportation connections.**Join us!** To apply, please share your **CV, motivation letter, and copies of degree certificates and academic transcripts** with us through our recruitment site (\"Apply now!” at the bottom of the page) **at the latest on 31st October 2026** 23.59pm (EET).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Postdoctoral Researcher in Speech and Language Technology
Postdoctoral Researcher in Speech and Language Technology

Aalto • Espoo

Hybrid
EUR 40,000 - 54,000
Doctoral Researcher in Speech and Language Technology
Doctoral Researcher in Speech and Language Technology

Aalto University • Finland

Hybrid
EUR 32,000 - 38,000
Doctoral Researcher in Speech and Language Technology
Doctoral Researcher in Speech and Language Technology

Aalto University • Espoo

Hybrid
EUR 32,000 - 38,000
Hybrid work model
International collaboration
Open science environment
+1
Postdoctoral Researcher in Speech and Language Technology
Postdoctoral Researcher in Speech and Language Technology

Aalto University • Finland

Hybrid
EUR 43,000 - 51,000
Postdoctoral Researcher in Speech and Language Technology
Postdoctoral Researcher in Speech and Language Technology

Aalto University • Kuhmo

Hybrid
EUR 43,000 - 51,000
Hybrid work model
International collaboration
Open science and publishing
Postdoctoral Researcher in Speech and Language Technology
Postdoctoral Researcher in Speech and Language Technology

Aalto University • Espoo

Hybrid
EUR 40,000 - 54,000
Hybrid model
Open science and open-source releases
Excellent computing infrastructure
PhD Researcher in Speech & Audio Watermarking & Detection
PhD Researcher in Speech & Audio Watermarking & Detection

Aalto University • Finland

Hybrid
EUR 32,000 - 38,000
Doctoral Researcher in Probabilistic Machine Learning
Doctoral Researcher in Probabilistic Machine Learning

Aalto • Espoo

Hybrid
EUR 36,000 - 40,000
Occupational health benefits
Social security in Finland
Doctoral Researcher in Backscatter-Assisted Wireless Positioning
Doctoral Researcher in Backscatter-Assisted Wireless Positioning

AALTO UNIVERSITY • Finland

Hybrid
EUR 35,000 - 39,000
Occupational healthcare
Employee benefits
Research Engineer
Research Engineer

Aalto • Espoo

Hybrid
EUR 41,000 - 51,000
Flexible working hours
Occupational health care
Supportive academic environment