Built independently by an author, for readers. Read the story and support ChapterPal

keyword

TTS augmentation

TTS augmentation is a machine learning technique in speech processing where text-to-speech synthesis systems are used to generate synthetic audio from text to expand training datasets. In automatic speech recognition and related audio tasks, models typically require large volumes of paired audio recordings and text transcriptions, which are often scarce or costly to collect for low-resource languages, dialects, or specialized vocabularies. By converting text-only sources into synthetic spoken utterances, TTS augmentation artificially produces labeled audio-text pairs. This approach increases the volume and diversity of training data, helping to improve the accuracy, generalization, and acoustic robustness of speech processing models without solely relying on manually recorded human speech.

1 item

Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation

Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation

Martijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky, Martijn Wieling

OrganizationsStanford UniversityUniversity of GroningenUniversity of Hawai’i at Mānoa

Why you should read this

Demonstrates how self-training and text-to-speech data augmentation can reduce word error rates by up to 25.5% when fine-tuning speech recognition models on as little as 24 minutes of transcribed audio across four diverse minority languages.

The performance of automatic speech recognition (ASR) systems has advanced substantially in recent years, particularly for languages for which a large amount of transcribed speech is available. Unfortunately, for low-resource languages, such as minority languages, regional languages or dialects, ASR performance generally remains much lower. In this study, we investigate whether data augmentation techniques could help improve low-resource ASR performance, focusing on four typologically diverse minority languages or language variants (West Germanic: Gronings, West-Frisian; Malayo-Polynesian: Besemah, Nasal). For all four languages, we examine the use of self-training, where an ASR system trained with the available human-transcribed data is used to generate transcriptions, which are then combined with the original data to train a new ASR system. For Gronings, for which there was a pre-existing text-to-speech (TTS) system available, we also examined the use of TTS to generate ASR training data from text-only sources. We find that using a self-training approach consistently yields improved performance (a relative WER reduction up to 20.5% compared to using an ASR system trained on 24 minutes of manually transcribed speech). The performance gain from TTS augmentation for Gronings was even stronger (up to 25.5% relative reduction in WER compared to a system based on 24 minutes of manually transcribed speech). In sum, our results show the benefit of using self-training or (if possible) TTS-generated data as an efficient solution to overcome the limitations of data availability for resource-scarce languages in order to improve ASR performance.

Added

2026-10-02