topic
reverberant environment
A reverberant environment is a space in which sound produced by a source reaches a listener not only directly but also through many time-delayed, attenuated copies created by reflections from surrounding surfaces such as walls, ceilings, and floors. These overlapping reflections persist after the source stops, forming a decaying tail of sound that distorts the signal by smearing temporal and spectral features, masking energy dips, and blurring speech intelligibility, in contrast to a discrete echo, which is a single distinguishable repetition. The severity of reverberation depends on the room's geometry, volume, and the absorptive properties of its surfaces, with hard, reflective materials producing longer decay and porous materials shortening it, and it is commonly quantified by the reverberation time, the interval for the sound level to drop 60 decibels after the source ceases. In computer science and audio engineering, reverberant environments are a central challenge for microphone arrays, beamforming, speech enhancement, echo suppression, blind source separation, and speech communication systems, since hands-free devices such as conference phones and smart speakers pick up substantial reflected energy, and accurate recording and reproduction of reverberant conditions is needed to test such devices, model room acoustics, and design algorithms that can compensate for or exploit the reflections.
2 items

Towards Robust Speech Representation Learning for Thousands of Languages
William Chen, Wangyou Zhang, Yifan Peng, Xinjian Li, Jinchuan Tian, Jiatong Shi, Xuankai Chang, Soumi Maiti, Karen Livescu, Shinji Watanabe
Why you should read this
Presents XEUS, a fully open self-supervised speech encoder pre-trained on over one million hours of audio across 4,057 languages that incorporates an acoustic dereverberation objective to achieve state-of-the-art multilingual speech recognition performance.
Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world’s 7000+ languages. We propose XEUS, a Cross-lingual Encoder for Universal Speech, trained on over 1 million hours of data across 4057 languages, extending the language coverage of SSL models 4-fold. We combine 1 million hours of speech from existing publicly accessible corpora with a newly created corpus of 7400+ hours from 4057 languages, which will be publicly released. To handle the diverse conditions of multilingual speech data, we augment the typical SSL masked prediction approach with a novel dereverberation objective, increasing robustness. We evaluate XEUS on several benchmarks, and show that it consistently outperforms or achieves comparable results to state-of-the-art (SOTA) SSL models across a variety of tasks. XEUS sets a new SOTA on the ML-SUPERB benchmark: it outperforms MMS 1B and w2v-BERT 2.0 v2 by 0.8% and 4.4% respectively, despite having less parameters or pre-training data. Checkpoints, code, and data are found in https://www.wavlab.org/activities/2024/xeus/.
Added
2026-10-03

GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration
Naoki Murata, Koichi Saito, Chieh-Hsin Lai, Yuhta Takida, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon
Why you should read this
Proposes a partially collapsed Gibbs sampling framework that enables pre-trained diffusion models to solve blind inverse problems like image deblurring and vocal dereverberation without requiring fine-tuning or specialized priors for the unknown measurement operator.
Pre-trained diffusion models have been successfully used as priors in a variety of linear inverse problems, where the goal is to reconstruct a signal from noisy linear measurements. However, existing approaches require knowledge of the linear operator. In this paper, we propose GibbsDDRM, an extension of Denoising Diffusion Restoration Models (DDRM) to a blind setting in which the linear measurement operator is unknown. GibbsDDRM constructs a joint distribution of the data, measurements, and linear operator by using a pre-trained diffusion model for the data prior, and it solves the problem by posterior sampling with an efficient variant of a Gibbs sampler. The proposed method is problem-agnostic, meaning that a pre-trained diffusion model can be applied to various inverse problems without fine-tuning. In experiments, it achieved high performance on both blind image deblurring and vocal dereverberation tasks, despite the use of simple generic priors for the underlying linear operators.
Added
2026-10-02
