Toward a realistic model of speech processing in the brain with self-supervised learning
Juliette MilletCharlotte CaucheteuxPierre OrhanYves BoubenecAlexandre GramfortEwan DunbarChristophe PallierJean-Remi King
Demonstrates that self-supervised neural networks trained on raw audio learn brain-like cortical hierarchies and functional specializations from realistic amounts of unlabeled speech, offering a biologically plausible computational framework for human language acquisition.
Recent advances in artificial intelligence have produced deep neural networks that generate internal representations similar to those found in the human brain. However, existing computational models of language remain biologically implausible because they rely on supervised labels, pre-tokenized text instead of raw sensory input, massive working memory windows, or unrealistic amounts of data—often equivalent to lifetimes of reading. Understanding how the human brain rapidly acquires speech processing capabilities with minimal, unlabeled sensory exposure is a critical open challenge in neuroscience and artificial intelligence.
The article evaluates whether self-supervised deep learning architectures trained directly on raw speech audio can account for the functional organization and behavioral patterns of speech processing in the human brain under biologically plausible data constraints.
To test this hypothesis, the researchers trained multiple variants of the wav2vec 2.0 neural network on 600 hours of unlabeled speech waveforms in English, French, and Mandarin, as well as on non-speech environmental sounds and fully supervised speech recognition. A 600-hour budget approximates the limited auditory exposure an infant receives during early language acquisition. Using linear encoding models with cross-validation, the authors compared internal network representations directly against whole-brain functional magnetic resonance imaging data collected from 412 adult native speakers who listened to naturalistic audiobooks. They also evaluated model representations against behavioral speech sound discrimination data from 386 human participants performing forced-choice perceptual tasks.
The analysis yielded four major findings. First, self-supervised training on 600 hours of raw audio enabled the model to significantly predict cortical responses to speech across the brain, performing modestly but significantly better than supervised models and substantially outperforming untrained models. Second, the functional hierarchy of the model's transformer layers mapped directly onto the anatomical hierarchy of human speech cortex, where early network layers best predicted low-level primary auditory cortices and deeper layers best predicted higher-order temporal and frontal areas. Third, the model developed speech- and language-specific representations matching those in human temporal cortex, with models trained on a participant's native language outperforming models trained on non-native speech, non-speech acoustic scenes, or random weights. Fourth, behavioral evaluations confirmed that self-supervised models mirrored human performance biases by discriminating native phonemes more accurately than non-native phonemes.
These findings indicate that explicit supervision and massive textual datasets are not required to reproduce brain-like speech representations. Simple self-supervised learning objectives applied to raw auditory waveforms provide a plausible computational framework for explaining how the brain organizes auditory input into hierarchical, language-specific structures. This challenges long-standing assumptions that the neural complexity of human language acquisition cannot be captured by concise algorithmic principles.
The article recommends that future research expand beyond adult fMRI by testing self-supervised architectures across broader language families, developmental cohorts of children, and high-temporal-resolution modalities such as electroencephalography or magnetoencephalography. It also suggests exploring architectures with realistic recurrent constraints and longer temporal receptive fields to bridge remaining gaps in semantic and syntactic processing.
Confidence in these findings is supported by the study's unusually large neuroimaging sample across three languages and consistent statistical significance across voxels. However, readers should consider important boundary conditions: functional magnetic resonance imaging provides limited temporal resolution, the model lacks human-like recurrent temporal constraints, and deep layers remain limited in their ability to capture rich semantics and complex syntax relative to text-based models, leaving the model at roughly 19% of the estimated noise ceiling across the whole brain.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Introduces wav2vec 2.0, the core self-supervised raw-waveform speech architecture whose representations and functional hierarchy are directly evaluated against human fMRI data in the source study.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). Presents HuBERT, a foundational masked self-supervised speech representation model that provides essential context for evaluating self-supervised acoustic models against biological speech perception.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Establishes the layer-wise probing methodology used to analyze hierarchical linguistic representation in deep neural networks, paralleling the cortical hierarchy mapping in the source paper.
- Paper: Beyond Mind-Reading: Multi-Voxel Pattern Analysis of fMRI Data, Kenneth A. Norman et al. (2006). Provides the foundational multi-voxel pattern analysis principles for mapping distributed neural activity and internal representations from fMRI brain recordings.
- Paper: Machine learning for neuroimaging with scikit-learn, Alexandre Abraham et al. (2014). Outlines standard machine learning encoding and decoding pipelines used to link high-dimensional computational model features to functional neuroimaging data.
- Paper: Robust Speech Recognition via Large-Scale Weak Supervision, Alec Radford et al. (2023). Scales speech representation learning to massive weakly supervised multilingual audio, extending beyond the self-supervised raw waveform models evaluated in the source.
- Paper: Unified Speech-Text Pre-training for Speech Translation and Recognition, Yun Tang et al. (2022). Extends self-supervised speech modeling by unifying raw speech and text representations in a shared pre-training framework for downstream speech tasks.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). Provides a comprehensive multi-task benchmark to evaluate advanced perception and reasoning across speech, music, and sound in audio-language models.
- Paper: Learn from your own latents and not from tokens: A sample-complexity theory, Daniel J. Korchinski et al. (2026). Develops a formal sample-complexity theory for self-supervised learning from internal latents rather than surface tokens, offering theoretical grounding for the data efficiency observed in the source.
