WAXAL: A large-scale multilingual African language speech corpus

Abdoulaye DiackPerry NelsonMohamedElfatih MohamedKhairSubhashini VenugopalanEmmanuel Asiedu BrempongTavonga SiyavoraBob MacDonaldUche OkonkwoSandy RitchieMandy Jordan

article2026arXiv13 citations

Presents WAXAL, an open-access corpus providing over 1,480 hours of speech recognition and text-to-speech data across 24 Sub-Saharan African languages spoken by more than 100 million people.

Listen

Voice-enabled technologies such as virtual assistants and automated transcription tools have expanded rapidly, yet their benefits remain concentrated in a small number of globally dominant languages. Over 100 million speakers across Sub-Saharan Africa face a major digital divide because high-quality, openly licensed speech corpora for their native languages remain scarce. This data gap is further compounded by regional linguistic complexities, including tonal variations, intricate morphology, and frequent code-switching, which hinder standard language modeling.

The article demonstrates the development and public release of WAXAL, a large-scale, multimodal speech corpus covering 24 Sub-Saharan African languages. Its primary objective is to evaluate, collect, and provide foundational datasets that enable the training and assessment of speech recognition and voice synthesis systems for critically underserved language communities.

To build the dataset between January 2021 and March 2024, researchers partnered with African academic institutions and community organizations across Ghana, Uganda, and Nigeria. For automated speech recognition, data collectors recorded spontaneous, natural speech from diverse speakers by prompting them with images across more than 50 topics in everyday indoor and outdoor environments. Local linguistic experts transcribed a subset of these recordings and screened them for quality and privacy. For text-to-speech synthesis, 72 contracted male and female voice actors recorded phonetically balanced scripts in controlled studio settings to ensure audio fidelity.

The initiative produced approximately 1,500 total hours of processed audio data. For speech recognition, the release delivers roughly 1,250 hours of transcribed natural speech comprising 224,767 instances across 14 languages, totaling 1.7 terabytes. For voice synthesis, the corpus provides approximately 235 hours of high-quality recordings across 13 languages, amounting to 99 gigabytes. Transcribers were compensated at above-average local wages, and all data was published openly under a permissive Creative Commons Attribution (CC-BY-4.0) license.

These findings and resources significantly lower the barrier to building inclusive language technologies, enabling both commercial and academic organizations to develop accessible applications without restrictive licensing barriers. By incorporating spontaneous speech and studio recordings from native speakers, the corpus provides a more representative baseline than previously available small-scale datasets, directly supporting digital preservation and regional digital equity.

Stakeholders and developers should leverage the openly accessible dataset to build and benchmark production-grade speech tools for the supported African languages. Future initiatives should focus on expanding transcription coverage beyond the initial 10% sample of collected audio and broadening sampling to capture wider dialectal and socio-linguistic variations. While manual screening mitigated the risk of inappropriate content, users should exercise appropriate oversight when deploying synthetic voices derived from the dataset.

arXiv: 2602.02734
  • Paper: Common Voice: A Massively-Multilingual Speech Corpus, Rosana Ardila et al. (2019). Read this earlier open multilingual speech-corpus release first to understand the corpus-building precedent against which WAXAL’s African-language coverage and collection approach are positioned.

No sufficiently relevant recommendations were found.

Cover for WAXAL: A large-scale multilingual African language speech corpus

Abstract

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at this https URL under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Data Collection Methodology
  • 3.1 ASR Data Collection and Annotation
  • 3.2 TTS Data Collection
  • 4 Dataset Statistics
  • 5 Limitations and Considerations
  • 5.1 Limitations
  • 5.2 Ethical Considerations
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — WAXAL provides paired ASR and TTS resources for African languages

    definition

    WAXAL is an openly released speech corpus covering 24 Sub-Saharan African languages spoken by more than 100 million people. It contains two complementary resources: transcribed, multi-speaker speech intended for automatic speech recognition (ASR), and high-quality, single-speaker recordings intended for text-to-speech (TTS). The released ASR component covers 14 languages and the TTS component 13; these counts are not disjoint. The corpus is available at https://huggingface.co/datasets/google/WaxalNLP under the CC-BY-4.0 license.

  2. Knowl 2 — ASR speech was elicited through images and recorded in natural environments

    experimental setup

    For WAXAL-ASR, participants described images in their native language rather than reading a script. The images covered at least 50 topics, with the goal of eliciting varied, relatively natural speech. Recordings were made in speakers’ natural environments and were at least 15 seconds long. Collection aimed for a broad age range and gender balance, and recorded speaker age, gender, language, and environment (for example, indoor, outdoor, or office).

  3. Knowl 3 — WAXAL-ASR transcription uses selective expert annotation and manual quality control

    model/method

    Local language experts transcribed a subset corresponding to 10% of the ASR audio collected by WAXAL’s partners. Each recording was transcribed by one annotator. Transcriptions used the language’s local script when available and otherwise used transliteration into the English alphabet. Audio and transcripts were checked for clarity, language accuracy, relevance to the image prompt, and appropriate content; personally identifiable information, if present, was removed during this review. The released ASR component contains approximately 1,250 hours of transcribed speech, so the corpus’s total collected ASR audio is not fully transcribed.

  4. Knowl 4 — WAXAL-TTS was recorded from community voice actors using balanced scripts

    experimental setup

    The WAXAL-TTS collection protocol used a phonetically balanced script of approximately 108,500 words for each of 10 target languages. Seventy-two community voice actors were contracted—36 men and 36 women. Recordings were made in a professional studio-like environment to obtain high-fidelity audio with minimal background noise, with a target of approximately 16 hours of clean, edited speech per actor. The released TTS component covers 13 languages and contains approximately 235 hours, as reported in the dataset statistics.

  5. Knowl 5 — Released ASR and TTS components differ in hours, instances, and storage size

    data/table

    WAXAL reports the following aggregate statistics for its released speech components. ASR has more instances and hours than TTS, while the full audio release occupies substantially more storage. The ASR and TTS language counts can overlap.

    Dataset Languages Total hours Instances
    WAXAL-ASR 14 ∼1,250\sim 1{,}250 224,767
    WAXAL-TTS 13 ∼235\sim 235 ∼17,660\sim 17{,}660

    The released ASR data occupies 1.7 TB, and the released TTS data occupies 99 GB.

  6. Knowl 6 — Transcribed ASR hours vary substantially across the 14 languages

    data/table

    The language-level chart reports approximately 1,244.4 transcribed ASR hours across the listed languages, consistent with the corpus’s rounded total of approximately 1,250 hours. Malagasy has the largest reported amount (182.5 hours); Acholi has the smallest (32.3 hours). The distribution indicates that the transcribed hours are not uniform across languages.

    Language Transcribed hours
    Malagasy 182.5
    Fulani 124.2
    Dagare 104.7
    Ikposo 103.8
    Akan 101.9
    Lingala 101.5
    Ewe 99.8
    Shona 99.2
    Dagbani 98.5
    Nyankole 50.9
    Soga 50.3
    Masaaba 48.8
    Luganda 46.0
    Acholi 32.3
  7. Knowl 7 — WAXAL-ASR has incomplete transcription and dialectal coverage

    limitation

    Only 10% of the ASR audio collected by WAXAL’s partners is transcribed in the current release. The dataset may not represent the full dialectal and sociolinguistic variation within each language, despite efforts to recruit diverse speakers. Because ASR speech is unscripted, some offensive or inappropriate content may occur; manual quality control was used to mitigate this risk. The diverse speakers and recording environments also make WAXAL-ASR unsuitable for training high-quality single-speaker TTS voices.

  8. Knowl 8 — ASR demographic metadata show many hours with unknown gender or age

    empirical result

    WAXAL’s ASR demographic charts report age-group recording hours of 4,591 for ages 18–30, 2,507 for ages 30–40, 1,166 for ages 40–50, 636 for ages 50 and above, and 2,137 with age unknown. The gender chart reports recording-hour shares of 22.9% male, 19.8% female, and 57.3% unknown. Thus, although collection aimed for gender balance and a wide age range, the released demographic summaries include substantial unknown information, especially for gender.

  9. Knowl 9 — WAXAL collection was carried out with African academic and community partners

    model/method

    The WAXAL data-collection effort ran from January 2021 through March 2024, with Google funding and technical mentorship and local collection work shared among African partners. Makerere University collected ASR data for five languages and TTS data for four; the University of Ghana collected ASR data for six languages and TTS data for two; Digital Umuganda collected ASR data for four languages; and Media Trust collected TTS data for four languages. These are the partners’ reported collection contributions, not a statement that each partner’s language set was distinct.

  10. Knowl 10 — WAXAL documents consent, privacy, and voice-use risks

    limitation

    ASR participants and TTS voice actors gave informed consent for recording and release for research under a permissive license. Personally identifiable information was removed from transcripts and metadata, but the authors note that a speaker’s language and accent may still allow ethnicity or race to be inferred. The authors also identify the possibility that released TTS voices could be used in unforeseen ways; they judge the potential community benefits of enabling speech technology for these languages to outweigh that risk. Local language experts hired for transcription were paid at rates described as well above the local hourly average.

Coverage note — Exact TTS hours by language are omitted because some language labels in the source chart are ambiguous; the overall TTS coverage, hours, and instance count are included. The image examples are illustrative rather than a separate methodological contribution.

References

  1. 1.Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber, “Common voice: A massively-multilingual speech corpus,” arXiv preprint arXiv:1912.06670, 2019.
  2. 2.Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 798–805.
  3. 3.Mark JF Gales, Kate M Knill, Anton Ragni, and Shakti P Rath, “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in Fourth International workshop on spoken language technologies for under-resourced languages (SLTU-2014). International Speech Communication Association (ISCA), 2014, pp. 16–23.
  4. 4.Alan W Black, “Cmu wilderness multilingual speech dataset,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 5971–5975.
  5. 5.Jorgen Valk and Tanel Alumäe, “Voxlingua107: a dataset for spoken language recognition,” in 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021, pp. 652–658.
  6. 6.Xinjian Li, Shinnosuke Takamichi, Takaaki Saeki, William Chen, Sayaka Shiota, and Shinji Watanabe, “Yodas: Youtube-oriented dataset for audio and speech,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–8.
  7. 7.Yifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui, Jinpeng Li, Bo Yang, Yexing Du, Ziyang Ma, Xunying Liu, Ziyuan Wang, et al., “Gigaspeech 2: An evolving, large-scale and multi-domain asr corpus for low-resource languages with automated crawling, transcription and refinement,” arXiv preprint arXiv:2406.11546, 2024.
  8. 8.Moussa Doumbouya, Lisa Einstein, and Chris Piech, “Using radio archives for low-resource speech recognition: Towards an intelligent virtual assistant for illiterate users,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2021, vol. 35.
  9. 9.Elodie Gauthier, Laurent Besacier, Sylvie Voisin, Michael Melese, and Uriel Pascal Elingui, “Collecting resources in sub-saharan african languages for automatic speech recognition: a case study of wolof,” in 10th Language Resources and Evaluation Conference (LREC 2016), 2016.
  10. 10.Daniel van Niekerk, Charl van Heerden, Marelie Davel, Neil Kleynhans, Oddur Kjartansson, Martin Jansche, and Linne Ha, “Rapid development of TTS corpora for four South African languages,” in Proc. Interspeech 2017, Stockholm, Sweden, Aug. 2017, pp. 2178–2182.
  11. 11.Josh Meyer, David Adelani, Edresson Casanova, Alp Oktem, Daniel Whitenack, Julian Weber, Salomon Kabongo Kabenamualu, Elizabeth Salesky, Iroro Orife, Colin Leong, Perez Ogayo, Chris Chinenye Emezue, Jonathan Mukiibi, Salomey Osei, Apelete Agbolo, Victor Akinode, Bernard Opoku, Olanrewaju Samuel, Jesujoba Alabi, and Shamsuddeen Hassan Muhammad, “Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus,” in Interspeech. 2022, ISCA.
  12. 12.VAANI Team, “Vaani: Capturing the language landscape for an inclusive digital india (phase 1),” https://vaani.iisc.ac.in/, 2025.

Citation

MLA
Diack, A., et al. “WAXAL: A Large-Scale Multilingual African Language Speech Corpus”. arXiv, 2026, http://arxiv.org/abs/2602.02734v3.
APA
Diack, A., Nelson, P., Agbesi, K., Nakalembe, A., MohamedKhair, M., Dube, V., Siyavora, T., Venugopalan, S., Hickey, J., Okonkwo, U., Bapna, A., Wiafe, I., Helegah, R. D., Atsakpo, E. D., Nutrokpor, C., Winful, F. B. P., Solaga, K. K., Abdulai, J.-D., Ekpezu, A. O., … Matias, Y. (2026). WAXAL: A Large-Scale Multilingual African Language Speech Corpus. arXiv. http://arxiv.org/abs/2602.02734v3
Chicago
Diack, A., P. Nelson, K. Agbesi, et al. 2026. “WAXAL: A Large-Scale Multilingual African Language Speech Corpus”. arXiv. http://arxiv.org/abs/2602.02734v3.
Harvard
Diack, A. et al. (2026) “WAXAL: A Large-Scale Multilingual African Language Speech Corpus”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.02734v3.
Vancouver
1. Diack A, Nelson P, Agbesi K, et al (2026) WAXAL: A Large-Scale Multilingual African Language Speech Corpus. arXiv

BibTeX

@article{diack2026waxal,
  title = {WAXAL: A Large-Scale Multilingual African Language Speech Corpus},
  author = {Diack, Abdoulaye and Nelson, Perry and Agbesi, Kwaku and Nakalembe, Angela and MohamedKhair, MohamedElfatih and Dube, Vusumuzi and Siyavora, Tavonga and Venugopalan, Subhashini and Hickey, Jason and Okonkwo, Uche and Bapna, Abhishek and Wiafe, Isaac and Helegah, Raynard Dodzi and Atsakpo, Elikem Doe and Nutrokpor, Charles and Winful, Fiifi Baffoe Payin and Solaga, Kafui Kwashie and Abdulai, Jamal-Deen and Ekpezu, Akon Obu and Niyonkuru, Audace and Rutunda, Samuel and Ishimwe, Boris and Melese, Michael and Bainomugisha, Engineer and Nakatumba-Nabende, Joyce and Katumba, Andrew and Babirye, Claire and Mukiibi, Jonathan and Kimani, Vincent and Kibacia, Samuel and Maina, James and Emmah, Fridah and Shekarau, Ahmed Ibrahim and Adamu, Ibrahim Shehu and Abdullahi, Yusuf and Lakougna, Howard and MacDonald, Bob and Shemtov, Hadar and Walcott-Bryant, Aisha and Cisse, Moustapha and Hassidim, Avinatan and Dean, Jeff and Matias, Yossi},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.02734v3},
  eprint = {2602.02734}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/