Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation
Martijn BarteldsNay SanBradley McDonnellDan JurafskyMartijn Wieling
Demonstrates how self-training and text-to-speech data augmentation can reduce word error rates by up to 25.5% when fine-tuning speech recognition models on as little as 24 minutes of transcribed audio across four diverse minority languages.
Recent advances in automated speech recognition have largely benefited major world languages, leaving regional, minority, and dialectal communities digitally underserved due to a lack of large transcribed audio datasets. Generating manual transcriptions for these languages requires substantial human effort and domain expertise, making data scarcity a central barrier to deploying language technologies in real-world documentation and preservation efforts.
The article demonstrates and evaluates whether practical data augmentation techniques can cost-effectively improve speech recognition performance for severely resource-scarce languages. Specifically, the authors assess the impact of self-training across four typologically diverse minority languages and examine synthetic speech generation using text-to-speech technology where existing systems allow.
To conduct this evaluation, the researchers analyzed four minority languages: Gronings, West-Frisian, Besemah, and Nasal. For each language, the total pool of human-transcribed speech was restricted to four hours. The team fine-tuned a standard multilingual pre-trained foundation model across varying training sizes (from 24 to 192 minutes) on a single high-performance graphics processing unit. They tested self-training by using an initial model to transcribe unannotated audio and retraining subsequent models on the combined data. Additionally, for Gronings, they augmented models using synthetic speech generated from text via an existing text-to-speech system.
The findings show that self-training consistently improves accuracy across all four languages, reducing word error rates by 6.3% to 13.9% relative to baseline models trained on just 24 minutes of transcribed audio. Iterative self-training with expanded unannotated audio delivered relative error reductions of up to 20.5% for Gronings, matching the performance of a model trained on double the volume of human-transcribed data. Furthermore, augmenting training with text-to-speech data yielded even larger improvements, achieving up to a 25.5% error reduction on out-of-domain evaluation and performing on par with models using four times the baseline amount of human transcriptions. In contrast, standard acoustic data modifications (such as altering pitch or adding noise) and computationally intensive continued pre-training yielded minimal or no meaningful gains.
These results demonstrate that organizations working on language preservation can significantly improve speech recognition systems without incurring the massive labor costs of manual transcription. Self-training leverages readily available unannotated speech recordings, while text-to-speech offers an effective shortcut whenever basic voice synthesizers and written text exist. This framework lowers the technological and financial barriers to building working speech tools for underrepresented communities.
For practical implementation, project teams should prioritize collecting unannotated audio for self-training or utilizing available text-to-speech systems over costly continued pre-training. However, decision-makers should note that increasing synthetic speech volume showed diminishing returns, likely due to single-speaker acoustic bias, and that acoustic performance remains lower than systems for high-resource languages. Future work should focus on multi-speaker synthetic generation, voice conversion methods, and testing broader sociolinguistic variations across larger speaker pools.
No sufficiently relevant recommendations were found.
- Paper: WST: Weakly Supervised Transducer for Automatic Speech Recognition, Dongji Gao et al. (2025). It advances low-resource self-training by learning robustly from noisy, weakly labeled transcripts, extending the source’s use of automatically generated labels.
- Paper: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale, Matthew Le et al. (2023). It extends synthetic speech augmentation by generating speech data for recognition within a broader multilingual speech-generation framework.
