Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Sub-Saharan African languages

Sub-Saharan African languages refers to the diverse group of indigenous languages spoken across the portion of the African continent situated south of the Sahara Desert. Representing thousands of distinct linguistic varieties across major families such as Niger-Congo, Nilo-Saharan, Afroasiatic, and Khoisan, these languages exhibit significant structural diversity, frequently featuring tone systems and complex noun classification systems. While spoken by hundreds of millions of people across highly multilingual societies, many of these languages are underrepresented in digital and computational technologies due to a historical scarcity of standardized orthographies, transcribed speech resources, and large-scale textual corpora.

1 item

WAXAL: A large-scale multilingual African language speech corpus

WAXAL: A large-scale multilingual African language speech corpus

Abdoulaye Diack, Perry Nelson, MohamedElfatih MohamedKhair, Subhashini Venugopalan, Emmanuel Asiedu Brempong, Tavonga Siyavora, Bob MacDonald, Uche Okonkwo, Sandy Ritchie, Mandy Jordan, Abhishek Bapna, Daan van Esch, Vusumuzi Dube, Jason Hickey, Ronit Levavi Morad, Yossi Matias, Jeff Dean, Aisha Walcott-Bryant, Avinatan Hassidim

OrganizationsAddis Ababa UniversityAfrican Institute for Mathematical Sciences SenegalBill & Melinda Gates FoundationDigital UmugandaGoogleLoud and Clear Comm. Ltd.Makerere UniversityMedia Trust Ltd.University of Ghana

Why you should read this

Presents WAXAL, an open-access corpus providing over 1,480 hours of speech recognition and text-to-speech data across 24 Sub-Saharan African languages spoken by more than 100 million people.

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at this https URL under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.

Added

2026-10-03