WAXAL: A large-scale multilingual African language speech corpus
Abdoulaye DiackPerry NelsonMohamedElfatih MohamedKhairSubhashini VenugopalanEmmanuel Asiedu BrempongTavonga SiyavoraBob MacDonaldUche OkonkwoSandy RitchieMandy Jordan
Presents WAXAL, an open-access corpus providing over 1,480 hours of speech recognition and text-to-speech data across 24 Sub-Saharan African languages spoken by more than 100 million people.
Voice-enabled technologies such as virtual assistants and automated transcription tools have expanded rapidly, yet their benefits remain concentrated in a small number of globally dominant languages. Over 100 million speakers across Sub-Saharan Africa face a major digital divide because high-quality, openly licensed speech corpora for their native languages remain scarce. This data gap is further compounded by regional linguistic complexities, including tonal variations, intricate morphology, and frequent code-switching, which hinder standard language modeling.
The article demonstrates the development and public release of WAXAL, a large-scale, multimodal speech corpus covering 24 Sub-Saharan African languages. Its primary objective is to evaluate, collect, and provide foundational datasets that enable the training and assessment of speech recognition and voice synthesis systems for critically underserved language communities.
To build the dataset between January 2021 and March 2024, researchers partnered with African academic institutions and community organizations across Ghana, Uganda, and Nigeria. For automated speech recognition, data collectors recorded spontaneous, natural speech from diverse speakers by prompting them with images across more than 50 topics in everyday indoor and outdoor environments. Local linguistic experts transcribed a subset of these recordings and screened them for quality and privacy. For text-to-speech synthesis, 72 contracted male and female voice actors recorded phonetically balanced scripts in controlled studio settings to ensure audio fidelity.
The initiative produced approximately 1,500 total hours of processed audio data. For speech recognition, the release delivers roughly 1,250 hours of transcribed natural speech comprising 224,767 instances across 14 languages, totaling 1.7 terabytes. For voice synthesis, the corpus provides approximately 235 hours of high-quality recordings across 13 languages, amounting to 99 gigabytes. Transcribers were compensated at above-average local wages, and all data was published openly under a permissive Creative Commons Attribution (CC-BY-4.0) license.
These findings and resources significantly lower the barrier to building inclusive language technologies, enabling both commercial and academic organizations to develop accessible applications without restrictive licensing barriers. By incorporating spontaneous speech and studio recordings from native speakers, the corpus provides a more representative baseline than previously available small-scale datasets, directly supporting digital preservation and regional digital equity.
Stakeholders and developers should leverage the openly accessible dataset to build and benchmark production-grade speech tools for the supported African languages. Future initiatives should focus on expanding transcription coverage beyond the initial 10% sample of collected audio and broadening sampling to capture wider dialectal and socio-linguistic variations. While manual screening mitigated the risk of inappropriate content, users should exercise appropriate oversight when deploying synthetic voices derived from the dataset.
- Paper: Common Voice: A Massively-Multilingual Speech Corpus, Rosana Ardila et al. (2019). Read this earlier open multilingual speech-corpus release first to understand the corpus-building precedent against which WAXAL’s African-language coverage and collection approach are positioned.
No sufficiently relevant recommendations were found.
