keyword
automatic speech recognition
Automatic speech recognition is a technology and subfield of artificial intelligence and computational linguistics that enables computer systems to identify, process, and convert human spoken language into written text. Frequently abbreviated as ASR and commonly referred to as speech-to-text, the technology processes acoustic audio signals by analyzing vocal frequencies, phonetic units, and linguistic context to generate accurate transcriptions of spoken utterances. Contemporary systems rely extensively on machine learning and deep neural network architectures, such as end-to-end neural models and recurrent transducers, to map raw audio features directly to text while accounting for variations in accents, background noise, dialects, and vocabulary. This enables a broad spectrum of real-world applications, including automated subtitling, voice-activated virtual assistants, transcription services, and hands-free human-computer interfaces.
4 items

WAXAL: A large-scale multilingual African language speech corpus
Abdoulaye Diack, Perry Nelson, MohamedElfatih MohamedKhair, Subhashini Venugopalan, Emmanuel Asiedu Brempong, Tavonga Siyavora, Bob MacDonald, Uche Okonkwo, Sandy Ritchie, Mandy Jordan, Abhishek Bapna, Daan van Esch, Vusumuzi Dube, Jason Hickey, Ronit Levavi Morad, Yossi Matias, Jeff Dean, Aisha Walcott-Bryant, Avinatan Hassidim
Why you should read this
Presents WAXAL, an open-access corpus providing over 1,480 hours of speech recognition and text-to-speech data across 24 Sub-Saharan African languages spoken by more than 100 million people.
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at this https URL under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.
Added
2026-10-03

AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Ke Li, Junteng Jia, Yuan Shangguan, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, Mike Seltzer
Why you should read this
Presents AudioChatLlama, an end-to-end framework that extends instruction-tuned language models to general-purpose speech reasoning and conversational tasks without requiring paired multimodal datasets by training an audio encoder to align speech prompts with text-based responses.
In this work, we extend the instruction-tuned Llama-2 model with end-to-end general-purpose speech processing and reasoning abilities while maintaining the wide range of original LLM capabilities, without using any carefully curated paired data. The resulting end-to-end model, named AudioChatLlama, can utilize audio prompts as a replacement for text and sustain a conversation. Such a model also has extended cross-modal capabilities such as being able to perform spoken question answering (QA), speech translation, and audio summarization amongst many other closed and open-domain tasks. This is unlike prior approaches in speech, in which LLMs are extended to handle audio for a limited number of pre-designated tasks. On both synthesized and recorded speech QA test sets, evaluations show that our end-to-end approach is on par with or outperforms cascaded systems (speech recognizer + LLM) in terms of modelling the response to a prompt. Furthermore, unlike cascades, our approach can interchange text and audio modalities and intrinsically utilize prior context in a conversation to provide better results.
Added
2026-10-02

WST: Weakly Supervised Transducer for Automatic Speech Recognition
Dongji Gao, Chenda Liao, Changliang Liu, Matthew Wiesner, Leibny Paola Garcia, Daniel Povey, Sanjeev Khudanpur, Jian Wu
Why you should read this
Proposes a weakly supervised transducer framework that trains end-to-end speech recognition models on transcripts with up to 70% error rates using a flexible training graph, outperforming existing temporal classification methods without requiring auxiliary models or confidence scoring.
The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available.
Added
2026-09-30

SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, Quoc V. Le
Why you should read this
Introduces an effective data augmentation method that applies time-frequency masking and warping directly to spectrogram features, enabling end-to-end speech recognition models to outperform complex hybrid systems on standard benchmarks.
We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consists of warping the features, masking blocks of frequency channels, and masking blocks of time steps. We apply SpecAugment on Listen, Attend and Spell networks for end-to-end speech recognition tasks. We achieve state-of-the-art performance on the LibriSpeech 960h and Swichboard 300h tasks, outperforming all prior work. On LibriSpeech, we achieve 6.8% WER on test-other without the use of a language model, and 5.8% WER with shallow fusion with a language model. This compares to the previous state-of-the-art hybrid system of 7.5% WER. For Switchboard, we achieve 7.2%/14.6% on the Switchboard/CallHome portion of the Hub5'00 test set without the use of a language model, and 6.8%/14.1% with shallow fusion, which compares to the previous state-of-the-art hybrid system at 8.3%/17.3% WER.
Added
2026-09-11
