Toward a realistic model of speech processing in the brain with self-supervised learning

Juliette MilletCharlotte CaucheteuxPierre OrhanYves BoubenecAlexandre GramfortEwan DunbarChristophe PallierJean-Remi King

article2022NeurIPS164 citations

Demonstrates that self-supervised neural networks trained on raw audio learn brain-like cortical hierarchies and functional specializations from realistic amounts of unlabeled speech, offering a biologically plausible computational framework for human language acquisition.

Listen

Recent advances in artificial intelligence have produced deep neural networks that generate internal representations similar to those found in the human brain. However, existing computational models of language remain biologically implausible because they rely on supervised labels, pre-tokenized text instead of raw sensory input, massive working memory windows, or unrealistic amounts of data—often equivalent to lifetimes of reading. Understanding how the human brain rapidly acquires speech processing capabilities with minimal, unlabeled sensory exposure is a critical open challenge in neuroscience and artificial intelligence.

The article evaluates whether self-supervised deep learning architectures trained directly on raw speech audio can account for the functional organization and behavioral patterns of speech processing in the human brain under biologically plausible data constraints.

To test this hypothesis, the researchers trained multiple variants of the wav2vec 2.0 neural network on 600 hours of unlabeled speech waveforms in English, French, and Mandarin, as well as on non-speech environmental sounds and fully supervised speech recognition. A 600-hour budget approximates the limited auditory exposure an infant receives during early language acquisition. Using linear encoding models with cross-validation, the authors compared internal network representations directly against whole-brain functional magnetic resonance imaging data collected from 412 adult native speakers who listened to naturalistic audiobooks. They also evaluated model representations against behavioral speech sound discrimination data from 386 human participants performing forced-choice perceptual tasks.

The analysis yielded four major findings. First, self-supervised training on 600 hours of raw audio enabled the model to significantly predict cortical responses to speech across the brain, performing modestly but significantly better than supervised models and substantially outperforming untrained models. Second, the functional hierarchy of the model's transformer layers mapped directly onto the anatomical hierarchy of human speech cortex, where early network layers best predicted low-level primary auditory cortices and deeper layers best predicted higher-order temporal and frontal areas. Third, the model developed speech- and language-specific representations matching those in human temporal cortex, with models trained on a participant's native language outperforming models trained on non-native speech, non-speech acoustic scenes, or random weights. Fourth, behavioral evaluations confirmed that self-supervised models mirrored human performance biases by discriminating native phonemes more accurately than non-native phonemes.

These findings indicate that explicit supervision and massive textual datasets are not required to reproduce brain-like speech representations. Simple self-supervised learning objectives applied to raw auditory waveforms provide a plausible computational framework for explaining how the brain organizes auditory input into hierarchical, language-specific structures. This challenges long-standing assumptions that the neural complexity of human language acquisition cannot be captured by concise algorithmic principles.

The article recommends that future research expand beyond adult fMRI by testing self-supervised architectures across broader language families, developmental cohorts of children, and high-temporal-resolution modalities such as electroencephalography or magnetoencephalography. It also suggests exploring architectures with realistic recurrent constraints and longer temporal receptive fields to bridge remaining gaps in semantic and syntactic processing.

Confidence in these findings is supported by the study's unusually large neuroimaging sample across three languages and consistent statistical significance across voxels. However, readers should consider important boundary conditions: functional magnetic resonance imaging provides limited temporal resolution, the model lacks human-like recurrent temporal constraints, and deep layers remain limited in their ability to capture rich semantics and complex syntax relative to text-based models, leaving the model at roughly 19% of the estimated noise ceiling across the whole brain.

arXiv: 2206.01685
Cover for Toward a realistic model of speech processing in the brain with self-supervised learning

Abstract

Several deep neural networks have recently been shown to generate activations similar to those of the brain in response to the same input. These algorithms, however, remain largely implausible: they require (1) extraordinarily large amounts of data, (2) unobtainable supervised labels, (3) textual rather than raw sensory input, and / or (4) implausibly large memory (e.g. thousands of contextual words). These elements highlight the need to identify algorithms that, under these limitations, would suffice to account for both behavioral and brain responses. Focusing on speech processing, we here hypothesize that self-supervised algorithms trained on the raw waveform constitute a promising candidate. Specifically, we compare a recent self-supervised model, wav2vec 2.0, to the brain activity of 412 English, French, and Mandarin individuals recorded with functional Magnetic Resonance Imaging (fMRI), while they listened to approximately one hour of audio books. First, we show that this algorithm learns brain-like representations with as little as 600 hours of unlabelled speech – a quantity comparable to what infants can be exposed to during language acquisition. Second, its functional hierarchy aligns with the cortical hierarchy of speech processing. Third, different training regimes reveal a functional specialization akin to the cortex: wav2vec 2.0 learns sound-generic, speech-specific and language-specific representations similar to those of the prefrontal and temporal cortices. Fourth, we confirm the similarity of this specialization with the behavior of 386 additional participants. These elements, resulting from the largest neuroimaging benchmark to date, show how self-supervised learning can account for a rich organization of speech processing in the brain, and thus delineate a path to identify the laws of language acquisition which shape the human brain.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Models
  • 2.1.1 Architecture
  • 2.1.2 Learning objective
  • 2.1.3 Training
  • 2.2 Functional MRI
  • 2.3 Brain score (R)
  • 2.4 Behavioral experiment
  • 3 Results
  • 4 Discussion
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — Hierarchical Mapping Between Transformer Layers and the Cortical Speech Processing Pathway

    empirical result

    When evaluating linear encoding models layer-by-layer across the 19 layers of wav2vec 2.0 (7 convolutional blocks and 12 transformer blocks), the model's depth hierarchy corresponds systematically to the anatomical hierarchy of human speech processing in functional Magnetic Resonance Imaging (fMRI):

    1. Acoustic / Early Auditory Cortex: The 7 convolutional feature encoder layers exhibit lower predictive brain scores than transformer layers. The early transformer layers (layers 1–4) best predict neural activity in primary and secondary auditory cortices (A1A1 and A2A2, located in Heschl's gyrus).
    2. Higher-Order Language Regions: Deeper transformer layers (layers 8–12) best predict neural responses in higher-level language-processing areas, specifically the superior temporal sulcus (STS), superior temporal gyrus (STG), and the inferior frontal gyrus (IFG / Broca's area).
    3. Motor Cortex Extension: This layer-to-cortex functional gradient extends into the supplementary motor and primary motor areas involved in laryngeal and articulatory mouth control across both hemispheres.
  2. Knowl 2 — Superiority of Self-Supervised Over Supervised Learning in Predicting Cortical Responses to Speech

    empirical result

    A wav2vec 2.0 model trained with a self-supervised contrastive and diversity objective on 600 hours of unlabelled speech predicts human fMRI responses across the cortex significantly better than the identical architecture trained with a supervised Connectionist Temporal Classification (CTC) loss for phoneme recognition on the same 600 hours of speech.

    Across 412 participants listening to naturalistic audiobooks, the self-supervised model achieves higher cross-validated Pearson correlation brain scores than the supervised model with an average voxel score improvement of ΔR=0.002\Delta R = 0.002 (p<10−6p < 10^{-6}). This demonstrates that explicit linguistic/phonetic supervision is not required to yield brain-like representations of spoken language from raw audio.

  3. Knowl 3 — Stepwise Cortical Specialization for Acoustic, Speech, and Native Language Representations

    empirical result

    Comparing fMRI encoding performance across different pretraining corpora reveals a hierarchical functional specialization in the human brain:

    1. Acoustic Learning: A wav2vec 2.0 model trained on 600 hours of non-speech acoustic scenes (Audioset) outperforms an untrained, randomly initialized model across cortical voxels (ΔR=0.006\Delta R = 0.006, p=10−31p = 10^{-31}).
    2. Speech Specificity: A model trained on 600 hours of non-native speech outperforms the non-speech acoustic model (ΔR=0.002\Delta R = 0.002, p=10−9p = 10^{-9}), showing that speech-specific representations emerge from self-supervised exposure to speech.
    3. Native Language Specificity: A model trained on 600 hours of the listener's native language achieves higher predictive scores than models trained on non-native languages (ΔR=0.002\Delta R = 0.002, p=10−15p = 10^{-15}).

    Voxels showing significant native-language specificity are predominantly concentrated in the superior temporal sulcus (STS) and middle temporal gyrus (MTG).

  4. Knowl 4 — Phonetic Discrimination Alignment Between Self-Supervised Models and Human Behavior

    empirical result

    In an ABX matching-to-sample phoneme discrimination task across ~6,000 sound triplets (508 English and 524 French phone pairs), self-supervised wav2vec 2.0 models reproduce human perceptual biases for native speech sounds:

    • Human Behavior (N=386N = 386): English and French native speakers discriminate native phone contrasts significantly more accurately than non-native contrasts (p<10−18p < 10^{-18}).
    • Model Behavior: Representations extracted from the most discriminative layer of wav2vec 2.0 (transformer layer 5, measured by Euclidean distance between representations) achieve significantly higher ABX discrimination accuracy on the language corresponding to their 600-hour pretraining corpus than on non-native speech (p<0.05p < 0.05).
    • Baselines: Randomly initialized models and models trained exclusively on non-speech environmental audio fail to develop this phonetic specialization and yield the lowest overall ABX discrimination accuracy.
  5. Knowl 5 — fMRI Linear Encoding Pipeline for Spoken Audio Neural Networks

    model/method

    To evaluate the similarity between internal neural network representations XX and voxel-wise fMRI blood-oxygen-level-dependent (BOLD) signals YY, a standard linear encoding framework is implemented:

    1. Feature Extraction: Latent representations are extracted from convolutional and transformer blocks using a sliding window of 10 seconds of raw speech waveform (stride =5= 5 seconds).
    2. Hemodynamic Response Alignment: Because model activations sample at 49–200 Hz while fMRI sampling is 0.5 Hz with delayed vascular dynamics, activations XtrainX_{\text{train}} are convolved with a canonical Glover hemodynamic response function (HRF) and downsampled to obtain temporally aligned features Xtrain′X'_{\text{train}}.
    3. Ridge Regression: A penalized linear mapping WW is fitted using ℓ2\ell_2-regularization (Ridge regression), where the hyperparameter λ∈[101,108]\lambda \in [10^1, 10^8] (20 log-spaced values) is optimized per voxel via nested cross-validation on the training set.
    4. Cross-Validation Evaluation: Using a 5-fold cross-validation scheme over time segments, brain similarity (the brain score RR) is computed as the Pearson correlation coefficient between true and predicted test BOLD responses: R=corr⁡(Ytest,WXtest′)R = \operatorname{corr}(Y_{\text{test}}, W X'_{\text{test}}), evaluated across subjects via a two-sided Wilcoxon signed-rank test with Benjamini–Hochberg False Discovery Rate (FDR) correction.
  6. Knowl 6 — Circular Mean Formulation for Identifying Optimal Predictive Model Layers Across Subjects

    equation

    To compute the representative optimal predictive layer k∗k^* for each cortical voxel across a cohort of subjects while avoiding regression-to-the-mean artifacts, the circular mean of individual optimal layer indices is defined as:

    k∗=angle⁡(1N∑s=1Nexp⁡(2iπksK+1))k^* = \operatorname{angle}\left(\frac{1}{N} \sum_{s=1}^{N} \exp\left(\frac{2i\pi k_s}{K + 1}\right)\right)

    where:

    • k∗k^* is the resulting circular mean optimal layer index mapped back onto the discrete layer space [1,K][1, K].
    • N=412N = 412 is the total number of human participants.
    • s∈{1,…,N}s \in \{1, \dots, N\} indexes the participant.
    • ks∈{1,…,K}k_s \in \{1, \dots, K\} is the layer index yielding the maximal Pearson correlation RR for participant ss at that specific voxel.
    • K=19K = 19 is the total number of model layers (7 convolutional feature encoder layers followed by 12 transformer blocks).
    • i=−1i = \sqrt{-1} is the imaginary unit.
    • angle⁡(z)=atan2⁡(Im⁡(z),Re⁡(z))\operatorname{angle}(z) = \operatorname{atan2}(\operatorname{Im}(z), \operatorname{Re}(z)) calculates the phase angle of the resulting complex sum.
  7. Knowl 7 — Architecture and 600-Hour Training Regimes for wav2vec 2.0 Variants

    experimental setup

    The wav2vec 2.0 framework used for brain encoding contains:

    • Feature Encoder: 7 temporal convolutional blocks (output dimension 512, temporal receptive field of 25 ms, stride of 20 ms, feature rate 49 Hz) that convert raw 16 kHz mono waveforms SS into latent representations zz.
    • Quantization Module: Discretizes zz into discrete latent sound vectors qq.
    • Context Network: 12 transformer blocks (model dimension 768, inner feed-forward dimension 3072, 8 attention heads) generating contextual representations cc.

    Training Regimes:

    1. Self-Supervised: Models are trained from scratch for 400k parameter updates using contrastive loss (predicting masked representations qq from cc) plus diversity loss on 600-hour subsets of French CommonVoice, English CommonVoice, MAGICDATA Mandarin, or filtered non-speech Audioset acoustic scenes.
    2. Supervised: The quantizer is replaced by a linear phoneme layer. All parameters are trained end-to-end for 400k updates using Connectionist Temporal Classification (CTC) loss over phonemized transcripts (32 French, 39 English, 33 Mandarin phonemes), reaching test Word Error Rates (WER) of 13.9% (French), 28.6% (English), and 4.6% (Mandarin).
  8. Knowl 8 — Multilingual Naturalistic Story-Listening fMRI Benchmark

    experimental setup

    The fMRI dataset aggregates 412 healthy native speakers passively listening to continuous spoken narratives in their native language (totaling 8.5 hours of distinct audio):

    1. Narratives Dataset (N=303N = 303): English-speaking participants listening to 15 un-scrambled stories ranging from 3 to 56 minutes (4.0 hours of unique audio, 36,018 total words, 4,004 unique words). Voxels are projected on the fsaverage6 surface space without spatial smoothing.
    2. The Little Prince Dataset (N=109N = 109): 48 English speakers (94 min audio), 33 Mandarin speakers (90 min audio), and 28 French speakers (97 min audio) listening to the audiobook read by native speakers over nine ~10-minute runs. Volumetric data are cortical-masked, projected onto fsaverage6 using FreeSurfer, and normalized to zero mean and unit variance per voxel, subject, and session.
  9. Knowl 9 — Representational Discrepancies and Noise-Ceiling Gaps Between wav2vec 2.0 and Cortical Speech Processing

    limitation

    Several functional and architectural discrepancies separate wav2vec 2.0 from human cortical speech processing:

    1. Temporal Non-Causality: Transformer self-attention layers attend bidirectionally across the entire 10-second input window, whereas biological auditory processing operates under causal, recurrent temporal constraints.
    2. Noise-Ceiling Gap: While wav2vec 2.0 activations explain up to 74% of the explainable variance (noise ceiling) in early auditory areas (Heschl's gyrus and sulcus), the model accounts for only ~19% of the noise ceiling on average across the entire cortex.
    3. Linguistic and Perceptual Fragility: The model lacks high-level recursive syntactic representations, captures substantially less semantic information than text-based language models, is excessively sensitive to acoustic band-pass filtering, and fails to display human-like categorical perception boundaries for phonemes.

Coverage note — None was omitted; all contributed models, encoding workflows, statistical equations, empirical fMRI and behavioral findings, and stated limitations are fully covered.

References

  1. 1.Alexandre Abraham, Fabian Pedregosa, Michael Eickenberg, Philippe Gervais, Andreas Mueller, Jean Kossaifi, Alexandre Gramfort, Bertrand Thirion, and Gaël Varoquaux. Machine learning for neuroimaging with scikit-learn. Frontiers in neuroinformatics, 8:14, 2014.
  2. 2.Richard Antonello, Javier S Turek, Vy Vo, and Alexander Huth. Low-dimensional structure in the space of language representations is reflected in brain responses. Advances in Neural Information Processing Systems, 34, 2021.
  3. 3.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. In LREC, 2020.
  4. 4.Alexei Baevski, H. Zhou, Abdel rahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. ArXiv, abs/2006.11477, 2020.
  5. 5.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906, 2021.
  6. 6.Gasper Begus, Alan Zhou, and Christina Zhao. Encoding of speech in convolutional layers and the brain stem based on language experience. bioRxiv, 2022.
  7. 7.Yoav Benjamini. Discovering the false discovery rate. Journal of the Royal Statistical Society: series B (statistical methodology), 72(4):405–416, 2010.
  8. 8.Julia Berezutskaya, Zachary V Freudenburg, Umut Güçlü, Marcel AJ van Gerven, and Nick F Ramsey. Neural tuning to low-level features of speech throughout the perisylvian cortex. Journal of Neuroscience, 37(33):7906–7920, 2017.
  9. 9.Ocke-Schwen Bohn. Cross-language and second language speech perception. The handbook of psycholinguistics, pages 213–239, 2017.
  10. 10.Lasse Borgholt, Jakob Drachmann Havtorn, Joakim Edin, Lars Maaløe, and Christian Igel. A brief overview of unsupervised neural speech representation learning. arXiv preprint arXiv:2203.01829, 2022.
  11. 11.Charlotte Caucheteux and Jean-Rémi King. Brains and algorithms partially converge in natural language processing. Communications Biology, 5(1):1–10, 2022.
  12. 12.Charlotte Caucheteux, Alexandre Gramfort, and Jean-Remi King. Disentangling syntax and semantics in the brain with deep networks. In Proceedings of the 38th International Conference on Machine Learning, pages 1336–1348. PMLR, July 2021a. ISSN: 2640-3498.
  13. 13.Charlotte Caucheteux, Alexandre Gramfort, and Jean-Remi King. Long-range and hierarchical language predictions in brains and algorithms. arXiv:2111.14232 [cs, q-bio], November 2021b. arXiv: 2111.14232.
  14. 14.Charlotte Caucheteux, Alexandre Gramfort, and Jean-Remi King. Model-based analysis of brain activity reveals the hierarchy of language in 305 subjects. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3635–3644, Punta Cana, Dominican Republic, November 2021c. Association for Computational Linguistics.
  15. 15.Charlotte Caucheteux, Alexandre Gramfort, and Jean-Rémi King. Deep language algorithms predict semantic comprehension from brain activity. Scientific Reports, 12(1):16327, September 2022. ISSN 2045-2322. doi: 10.1038/s41598-022-20460-9. Number: 1 Publisher: Nature Publishing Group.
  16. 16.Noam Chomsky. Linguistics and brain science. Image, language, brain, pages 13–28, 2000.
  17. 17.Radoslaw M Cichy and Daniel Kaiser. Deep neural networks as scientific models. Trends in cognitive sciences, 23(4):305–317, 2019.
  18. 18.Beijing Magic Data Technology Co. Magic data chinese mandarin conversational speech. http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101, 2019.
  19. 19.Christophe Destrieux, Bruce Fischl, Anders Dale, and Eric Halgren. Automatic parcellation of human cortical gyri and sulci using standard anatomical nomenclature. NeuroImage, 53(1):1–15, October 2010. ISSN 1053-8119. doi: 10.1016/j.neuroimage.2010.06.010.
  20. 20.Benjamin K Dichter, Jonathan D Breshears, Matthew K Leonard, and Edward F Chang. The control of vocal pitch in human laryngeal motor cortex. Cell, 174(1):21–31, 2018.
  21. 21.Emmanuel Dupoux. Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner. Cognition, 173:43–59, 2018.
  22. 22.Bruce Fischl. Freesurfer. Neuroimage, 62(2):774–781, 2012.
  23. 23.Angela D Friederici. The neurobiology of language comprehension. In Language comprehension: A biological perspective, pages 265–304. Springer, 1999.
  24. 24.Jack Gallant. in ’reading minds’. Nature, 502(7472):428, 2013.
  25. 25.Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 776–780. IEEE, 2017.
  26. 26.Jill Gilkerson, Jeffrey A Richards, Steven F Warren, Judith K Montgomery, Charles R Greenwood, D Kimbrough Oller, John HL Hansen, and Terrance D Paul. Mapping the early language environment using all-day recordings and automated analysis. American journal of speech-language pathology, 26(2):248–265, 2017.
  27. 27.Ariel Goldstein, Zaid Zada, Eliav Buchnik, Mariano Schain, Amy Price, Bobbi Aubrey, Samuel A Nastase, Amir Feder, Dotan Emanuel, Alon Cohen, et al. Shared computational principles for language processing in humans and deep language models. Nature neuroscience, 25(3):369–380, 2022.
  28. 28.Alex Graves. Connectionist temporal classification. In Supervised Sequence Labelling with Recurrent Neural Networks, pages 61–93. Springer, 2012.
  29. 29.Peter Hagoort. On broca, brain, and binding: a new framework. Trends in cognitive sciences, 9(9):416–423, 2005.
  30. 30.Betty Hart and Todd R Risley. American parenting of language-learning children: Persisting differences in family-child interactions observed in natural home environments. Developmental psychology, 28(6):1096, 1992.
  31. 31.Joseph Henrich, Steven J. Heine, and Ara Norenzayan. The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3):61–83, June 2010. ISSN 0140-525X, 1469-1825. doi: 10.1017/S0140525X0999152X.
  32. 32.Gregory Hickok and David Poeppel. The cortical organization of speech processing. Nature reviews neuroscience, 8(5):393–402, 2007.
  33. 33.Nicholas Huang, Malcolm Slaney, and Mounya Elhilali. Connecting deep neural networks to physical, perceptual, and electrophysiological auditory signals. Frontiers in neuroscience, 12:532, 2018.
  34. 34.L. G. Humphreys. Acquisition and extinction of verbal expectations in a situation analogous to conditioning. Journal of Experimental Psychology, 25:294–301, 1939. ISSN 0022-1015. doi: 10.1037/h0053555. Place: US Publisher: American Psychological Association.
  35. 35.Alexander G. Huth, Wendy A. de Heer, Thomas L. Griffiths, Frédéric E. Theunissen, and Jack L. Gallant. Natural speech reveals the semantic maps that tile human cerebral cortex. Nature, 532(7600):453–458, April 2016. ISSN 0028-0836, 1476-4687. doi: 10.1038/nature17637.
  36. 36.Shailee Jain and Alexander Huth. Incorporating context into language encoding models for fmri. Advances in neural information processing systems, 31, 2018.
  37. 37.Alexander JE Kell and Josh H McDermott. Deep neural network models of sensory systems: windows onto the role of task constraints. Current opinion in neurobiology, 55:121–132, 2019.
  38. 38.Alexander JE Kell, Daniel LK Yamins, Erica N Shook, Sam V Norman-Haignere, and Josh H McDermott. A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy. Neuron, 98(3):630–644, 2018.
  39. 39.Spencer Kellis, Kai Miller, Kyle Thomson, Richard Brown, Paul House, and Bradley Greger. Decoding spoken words using local field potentials recorded from the cortical surface. Journal of neural engineering, 7(5):056007, 2010.
  40. 40.Tim C Kietzmann, Patrick McClure, and Nikolaus Kriegeskorte. Deep neural networks in computational neuroscience. BioRxiv, page 133504, 2018.
  41. 41.Takuya Koumura, Hiroki Terashima, and Shigeto Furukawa. Cascaded tuning to amplitude modulation for natural sound recognition. Journal of Neuroscience, 39(28):5517–5533, 2019.
  42. 42.Nikolaus Kriegeskorte. Deep neural networks: a new framework for modeling biological vision and brain information processing. Annual review of vision science, 1:417–446, 2015.
  43. 43.Patricia K Kuhl, Barbara T Conboy, Denise Padden, Tobey Nelson, and Jessica Pruitt. Early speech perception and later language development: Implications for the" critical period". Language learning and development, 1(3-4):237–264, 2005.
  44. 44.Yair Lakretz, Théo Desbordes, Jean-Rémi King, Benoît Crabbé, Maxime Oquab, and Stanislas Dehaene. Can RNNs learn Recursive Nested Subject-Verb Agreements? arXiv:2101.02258 [cs], January 2021. URL http://arxiv.org/abs/2101.02258. arXiv: 2101.02258.
  45. 45.Y. Lerner, C. J. Honey, L. J. Silbert, and U. Hasson. Topographic Mapping of a Hierarchy of Temporal Receptive Windows Using a Narrated Story. Journal of Neuroscience, 31(8):2906–2915, February 2011. ISSN 0270-6474, 1529-2401. doi: 10.1523/JNEUROSCI.3684-10.2011.
  46. 46.Jixing Li, Shohini Bhattasali, Shulin Zhang, Berta Franzluebbers, Wen-Ming Luh, R Nathan Spreng, Jonathan R Brennan, Yiming Yang, Christophe Pallier, and John T Hale. Le petit prince: A multilingual fmri corpus using ecological stimuli. bioRxiv, 2021.
  47. 47.Saima Malik-Moraleda, Dima Ayyash, Jeanne Gallée, Josef Affourtit, Malte Hoffmann, Zachary Mineroff, Olessia Jouravlev, and Evelina Fedorenko. An investigation across 45 languages and 12 language families reveals a universal language network. Nature Neuroscience, 25(8):1014–1019, August 2022. ISSN 1546-1726. doi: 10.1038/s41593-022-01114-5. Number: 8 Publisher: Nature Publishing Group.
  48. 48.Nima Mesgarani, Connie Cheung, Keith Johnson, and Edward F. Chang. Phonetic Feature Encoding in Human Superior Temporal Gyrus. Science, 343(6174):1006–1010, February 2014. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.1245994. Publisher: American Association for the Advancement of Science Section: Report.
  49. 49.Juliette Millet and Ewan Dunbar. Do self-supervised speech models develop human-like perception biases? AAAI 2022 Workshop, Self-Supervised Learning for Audio and Speech Processing, 2022.
  50. 50.Juliette Millet and Jean-Remi King. Inductive biases, pretraining and fine-tuning jointly account for brain responses to speech. arXiv preprint arXiv:2103.01032, 2021.
  51. 51.Juliette Millet, Ioana Chitoran, and Ewan Dunbar. Predicting non-native speech perception using the perceptual assimilation model and state-of-the-art acoustic models. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 661–673, Online, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.conll-1.51.
  52. 52.Tom M Mitchell, Svetlana V Shinkareva, Andrew Carlson, Kai-Min Chang, Vicente L Malave, Robert A Mason, and Marcel Adam Just. Predicting human brain activity associated with the meanings of nouns. science, 320(5880):1191–1195, 2008.
  53. 53.Emily M Mugler, James L Patton, Robert D Flint, Zachary A Wright, Stephan U Schuele, Joshua Rosenow, Jerry J Shih, Dean J Krusienski, and Marc W Slutzky. Direct classification of all american english phonemes using signals from functional speech motor cortex. Journal of neural engineering, 11(3):035015, 2014.
  54. 54.Thomas Naselaris, Kendrick N Kay, Shinji Nishimoto, and Jack L Gallant. Encoding and decoding in fmri. Neuroimage, 56(2):400–410, 2011.
  55. 55.Samuel A Nastase, Yun-Fei Liu, Hanna Hillman, Asieh Zadbood, Liat Hasenfratz, Neggin Keshavarzian, Janice Chen, Christopher J Honey, Yaara Yeshurun, Mor Regev, et al. Narratives: fmri data for evaluating models of naturalistic language comprehension. preprint. Neuroscience, December, pages 2020–06, 2020.
  56. 56.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210. IEEE, 2015.
  57. 57.Ankita Pasad, Ju-Chieh Chou, and Karen Livescu. Layer-wise analysis of a self-supervised speech representation model. arXiv preprint arXiv:2107.04734, 2021.
  58. 58.Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830, 2011.
  59. 59.Francisco Pereira, Bin Lou, Brianna Pritchett, Samuel Ritter, Samuel J Gershman, Nancy Kanwisher, Matthew Botvinick, and Evelina Fedorenko. Toward a universal decoder of linguistic meaning from brain activation. Nature communications, 9(1):1–13, 2018.
  60. 60.Christopher I. Petkov, Yukiko Kikuchi, Alice E. Milne, Mortimer Mishkin, Josef P. Rauschecker, and Nikos K. Logothetis. Different forms of effective connectivity in primate frontotemporal pathways. Nature Communications, 6(1):6000, January 2015. ISSN 2041-1723. doi: 10.1038/ncomms7000. Number: 1 Publisher: Nature Publishing Group.
  61. 61.David Poeppel, Karen Emmorey, Gregory Hickok, and Liina Pylkkänen. Towards a new neurobiology of language. Journal of Neuroscience, 32(41):14125–14131, 2012.
  62. 62.Peng Qian, Xipeng Qiu, and Xuanjing Huang. Bridging lstm architecture and the neural dynamics during reading. arXiv preprint arXiv:1604.06635, 2016.
  63. 63.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  64. 64.Thomas Schatz. ABX-discriminability measures and applications. PhD thesis, Université Paris 6 (UPMC), 2016.
  65. 65.Martin Schrimpf, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A Hosseini, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences, 118(45), 2021.
  66. 66.Dan Schwartz, Mariya Toneva, and Leila Wehbe. Inducing brain-relevant bias in natural language processing models. Advances in neural information processing systems, 32, 2019.
  67. 67.Shihab Shamma, Prachi Patel, Shoutik Mukherjee, Guilhem Marion, Bahar Khalighinejad, Cong Han, Jose Herrero, Stephan Bickel, Ashesh Mehta, and Nima Mesgarani. Learning speech production and perception through sensorimotor interactions. Cerebral cortex communications, 2(1):tgaa091, 2021.
  68. 68.Cory Stephenson, Jenelle Feather, Suchismita Padhy, Oguz Elibol, Hanlin Tang, Josh McDermott, and SueYeon Chung. Untangling in invariant speech recognition. Advances in neural information processing systems, 32, 2019.
  69. 69.Jessica AF Thompson, Yoshua Bengio, Elia Formisano, and Marc Schönwiesner. Training neural networks to recognize speech increased their correspondence to the human auditory pathway but did not yield a shared hierarchy of acoustic features. bioRxiv, 2021.
  70. 70.Mariya Toneva and Leila Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). arXiv:1905.11833 [cs, q-bio], November 2019. arXiv: 1905.11833.
  71. 71.Aditya R Vaidya, Shailee Jain, and Alexander G Huth. Self-supervised models of audio effectively explain human cortical responses to speech. arXiv preprint arXiv:2205.14252, 2022.
  72. 72.Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. arXiv preprint arXiv:2101.00390, 2021.
  73. 73.Lotte Weerts, Stuart Rosen, Claudia Clopath, and Dan FM Goodman. The psychometrics of automatic speech recognition. bioRxiv, 2021.
  74. 74.Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. Self-training and Pre-training are Complementary for Speech Recognition, October 2020. arXiv:2010.11430 [cs, eess].
  75. 75.Daniel LK Yamins and James J DiCarlo. Using goal-driven deep learning models to understand sensory cortex. Nature neuroscience, 19(3):356–365, 2016.

Citation

MLA
MILLET, J., et al. “Toward a Realistic Model of Speech Processing in the Brain with Self-supervised Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 33428–43, https://proceedings.neurips.cc/paper_files/paper/2022/file/d81ecfc8fb18e833a3fa0a35d92532b8-Paper-Conference.pdf.
APA
MILLET, J., Caucheteux, C., orhan, . pierre ., Boubenec, Y., Gramfort, A., Dunbar, E., Pallier, C., & King, J.-R. (2022). Toward a realistic model of speech processing in the brain with self-supervised learning. Advances in Neural Information Processing Systems, 35, 33428–33443. https://proceedings.neurips.cc/paper_files/paper/2022/file/d81ecfc8fb18e833a3fa0a35d92532b8-Paper-Conference.pdf
Chicago
MILLET, J., C. Caucheteux, . pierre . orhan, et al. 2022. “Toward a Realistic Model of Speech Processing in the Brain with Self-supervised Learning”. Advances in Neural Information Processing Systems 35: 33428–43. https://proceedings.neurips.cc/paper_files/paper/2022/file/d81ecfc8fb18e833a3fa0a35d92532b8-Paper-Conference.pdf.
Harvard
MILLET, J. et al. (2022) “Toward a realistic model of speech processing in the brain with self-supervised learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 33428–33443. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d81ecfc8fb18e833a3fa0a35d92532b8-Paper-Conference.pdf.
Vancouver
1. MILLET J, Caucheteux C, orhan pierre, Boubenec Y, Gramfort A, Dunbar E, Pallier C, King J-R (2022) Toward a realistic model of speech processing in the brain with self-supervised learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 33428–33443

BibTeX

@inproceedings{millet2022toward,
  title = {Toward a realistic model of speech processing in the brain with self-supervised learning},
  author = {MILLET, Juliette and Caucheteux, Charlotte and orhan, pierre and Boubenec, Yves and Gramfort, Alexandre and Dunbar, Ewan and Pallier, Christophe and King, Jean-Remi},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {33428-33443},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d81ecfc8fb18e833a3fa0a35d92532b8-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors