Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech
Aditya R. VaidyaShailee JainAlexander Huth
Demonstrates that intermediate representations from self-supervised speech models accurately predict fMRI responses in the human auditory cortex and mirror the brain's hierarchical processing from acoustic to semantic features better than traditional acoustic and supervised baselines.
Understanding how the human brain processes spoken language is fundamental to neuroscience and artificial intelligence, yet models of human auditory processing have historically relied on hand-crafted acoustic filters or fully supervised speech recognition systems. These traditional baselines often fail to capture the full hierarchy of speech comprehension in continuous, real-world listening conditions. The article evaluates whether self-supervised speech representation learning—a machine learning paradigm that learns statistical regularities directly from raw audio without human annotations—can more accurately explain and predict human cortical responses to natural speech.
To test this, the authors built linearized voxel-wise encoding models using functional magnetic resonance imaging data collected from seven healthy adult participants who listened to over five hours of natural narrative stories. The researchers evaluated representations from every layer across four self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) alongside a supervised speech recognition model (Deep Speech 2), hand-engineered acoustic baselines (spectrotemporal modulations and filter banks), mid-level phonemic features, and high-level lexical and language models (word embeddings and GPT). Linear ridge regression was used to predict cortical blood-oxygen-level-dependent responses, and statistical probes were used to inspect what acoustic and linguistic properties were captured across network layers.
The investigation produced four primary findings. First, self-supervised audio representations—particularly the upper-middle layers of HuBERT and wav2vec 2.0—significantly outperformed hand-engineered acoustic baselines and supervised neural network features at predicting cortical responses across the entire cortex and within the auditory cortex. Second, encoding performance varied systematically by depth: early network layers best predicted low-level sensory regions such as primary auditory cortex, whereas upper-middle layers best predicted higher-level semantic regions like the angular gyrus and precuneus. Third, variance partitioning confirmed that the best-performing layer (layer 9 of HuBERT) completely encompassed the predictive variance of low-level spectrotemporal and mid-level phoneme features while substantially overlapping with high-level word embeddings. Fourth, direct linear probing confirmed that self-supervised architectures spontaneously develop an internal representational hierarchy—moving from raw spectral features in early layers to phonetic and lexical structures in deeper layers—mirroring human cortical organization without ever being trained on transcripts or labels.
These findings demonstrate that self-supervised audio models automatically learn rich, multi-tiered linguistic representations that transfer remarkably well to biological systems. In practical terms, this establishes self-supervised learning as a superior computational foundation for modeling auditory neuroscience, offering research and development teams higher modeling fidelity without the substantial data-labeling costs and annotation overhead associated with supervised systems. Although deep self-supervised audio representations approach the predictive power of static word embeddings in auditory regions, a clear performance gap remains between pure audio models and dedicated contextual text models like GPT in higher-level semantic areas, confirming that audio-only representations do not fully replace word-level language models.
Future research and development efforts should focus on integrating self-supervised speech representations into end-to-end neurocomputational pipelines, exploring multimodal models that unify speech audio with high-level contextual text transformers, and validating these models across diverse languages and non-narrative audio stimuli. The primary limitations of the study include its reliance on a small sample size (seven subjects) listening exclusively to English-language narratives and the intrinsic temporal smoothing of functional magnetic resonance imaging. Nevertheless, because the predictive performance gains were statistically significant and consistent across all evaluated participants and regions of interest, confidence in the primary conclusion—that self-supervised speech models effectively capture the cortical hierarchy of speech processing—remains high.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). Read HuBERT first to understand the masked-unit prediction model whose layer representations the study tests against cortical responses.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Its wav2vec 2.0 architecture and self-supervised training objective clarify one of the study’s strongest-performing speech representations.
No sufficiently relevant recommendations were found.
