Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech

Aditya R. VaidyaShailee JainAlexander Huth

article2022ICML98 citations

Demonstrates that intermediate representations from self-supervised speech models accurately predict fMRI responses in the human auditory cortex and mirror the brain's hierarchical processing from acoustic to semantic features better than traditional acoustic and supervised baselines.

Listen

Understanding how the human brain processes spoken language is fundamental to neuroscience and artificial intelligence, yet models of human auditory processing have historically relied on hand-crafted acoustic filters or fully supervised speech recognition systems. These traditional baselines often fail to capture the full hierarchy of speech comprehension in continuous, real-world listening conditions. The article evaluates whether self-supervised speech representation learning—a machine learning paradigm that learns statistical regularities directly from raw audio without human annotations—can more accurately explain and predict human cortical responses to natural speech.

To test this, the authors built linearized voxel-wise encoding models using functional magnetic resonance imaging data collected from seven healthy adult participants who listened to over five hours of natural narrative stories. The researchers evaluated representations from every layer across four self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) alongside a supervised speech recognition model (Deep Speech 2), hand-engineered acoustic baselines (spectrotemporal modulations and filter banks), mid-level phonemic features, and high-level lexical and language models (word embeddings and GPT). Linear ridge regression was used to predict cortical blood-oxygen-level-dependent responses, and statistical probes were used to inspect what acoustic and linguistic properties were captured across network layers.

The investigation produced four primary findings. First, self-supervised audio representations—particularly the upper-middle layers of HuBERT and wav2vec 2.0—significantly outperformed hand-engineered acoustic baselines and supervised neural network features at predicting cortical responses across the entire cortex and within the auditory cortex. Second, encoding performance varied systematically by depth: early network layers best predicted low-level sensory regions such as primary auditory cortex, whereas upper-middle layers best predicted higher-level semantic regions like the angular gyrus and precuneus. Third, variance partitioning confirmed that the best-performing layer (layer 9 of HuBERT) completely encompassed the predictive variance of low-level spectrotemporal and mid-level phoneme features while substantially overlapping with high-level word embeddings. Fourth, direct linear probing confirmed that self-supervised architectures spontaneously develop an internal representational hierarchy—moving from raw spectral features in early layers to phonetic and lexical structures in deeper layers—mirroring human cortical organization without ever being trained on transcripts or labels.

These findings demonstrate that self-supervised audio models automatically learn rich, multi-tiered linguistic representations that transfer remarkably well to biological systems. In practical terms, this establishes self-supervised learning as a superior computational foundation for modeling auditory neuroscience, offering research and development teams higher modeling fidelity without the substantial data-labeling costs and annotation overhead associated with supervised systems. Although deep self-supervised audio representations approach the predictive power of static word embeddings in auditory regions, a clear performance gap remains between pure audio models and dedicated contextual text models like GPT in higher-level semantic areas, confirming that audio-only representations do not fully replace word-level language models.

Future research and development efforts should focus on integrating self-supervised speech representations into end-to-end neurocomputational pipelines, exploring multimodal models that unify speech audio with high-level contextual text transformers, and validating these models across diverse languages and non-narrative audio stimuli. The primary limitations of the study include its reliance on a small sample size (seven subjects) listening exclusively to English-language narratives and the intrinsic temporal smoothing of functional magnetic resonance imaging. Nevertheless, because the predictive performance gains were statistically significant and consistent across all evaluated participants and regions of interest, confidence in the primary conclusion—that self-supervised speech models effectively capture the cortical hierarchy of speech processing—remains high.

Vaidya et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech

Abstract

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models’ ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.

Table of Contents

  • 1. Introduction
  • 2. Natural Language fMRI Experiment
  • 3. Voxel-wise Encoding Models
  • 3.1. Extracting Speech Features from SSL Models
  • 3.2. Supervised Speech Model Baseline
  • 3.3. Baselines
  • 4. Experiments
  • 4.1. Encoding Performance Comparisons
  • 4.2. Voxel-wise Layer Selectivity
  • 4.3. Partitioning Explained Variance between SSL Models and Hand-Engineered Baselines
  • 4.4. Probing SSL Models for Linguistic Structure
  • 5. Conclusion
  • Acknowledgments
  • References
  • A. Voxel-wise encoding models of speech
  • A.1. Performance in speech ROIs
  • B. Voxel-wise layer selectivity
  • C. Probing SSL model representations for linguistic structure
  • D. MRI acquisition, preprocessing, and experiment details
  • D.1. Participants
  • D.2. Stimulus preparation and presentation
  • D.3. Acquisition parameters
  • D.4. Processing MRI data
  • D.5. Defining regions of interest (ROIs)

Knowls

  1. Knowl 1 — Self-supervised speech features improve cortical response prediction

    empirical result

    For fMRI responses to natural English narratives, encoding models built from self-supervised speech model features predicted cortical activity better than models using hand-engineered acoustic features or phoneme articulations. Across the whole cortex, the best layers of the self-supervised models and supervised Deep Speech 2 exceeded the spectrotemporal and articulatory baselines; the best layers of wav2vec 2.0 and HuBERT also exceeded Deep Speech 2, despite the models being trained on the same amount of speech data. HuBERT layer 9 was the best-performing representation overall and approached the performance of word embeddings, although GPT features remained substantially better. Within broadly defined auditory cortex, most self-supervised and Deep Speech 2 layers exceeded the hand-constructed feature spaces, including word embeddings.

  2. Knowl 2 — Encoding performance peaks in intermediate speech-model layers

    empirical result

    Across APC, wav2vec, wav2vec 2.0, HuBERT, and Deep Speech 2, cortical encoding performance generally rose from the shallow layers toward the upper-middle layers and then declined in the deepest layers. This pattern was clearest for wav2vec 2.0, HuBERT, and Deep Speech 2. In auditory cortex, lower layers could remain comparatively effective in low-level auditory regions, while upper-middle layers performed best more broadly. Thus, the layer that best predicts a response depends on both representation depth and the cortical region being predicted.

  3. Knowl 3 — HuBERT layer preference varies along the cortical processing hierarchy

    empirical result

    For each participant, the authors applied PCA to a matrix whose rows were voxels and whose columns were HuBERT layers, with each entry giving that voxel's encoding performance for that layer; rows and columns were centered before PCA. The first principal component explained 54% of the variance in layer-wise performance, while every later component explained less than 10%. Its loadings distinguished lower HuBERT layers from upper-middle layers: primary auditory cortex voxels tended to prefer lower layers, whereas angular gyrus and precuneus voxels tended to prefer upper-middle layers. The first-component score was positively correlated with word-embedding encoding performance across cortex (ρ=0.449\rho=0.449) and negatively correlated with spectrotemporal-feature performance within auditory cortex (ρ=−0.330\rho=-0.330), consistent with later-layer preference tracking higher-level information.

  4. Knowl 4 — HuBERT's best layer overlaps acoustic, phonetic, and lexical information

    empirical result

    Pairwise variance partitioning compared HuBERT layers 1 and 9 with spectrotemporal features, 14-feature phoneme articulations, and word embeddings. HuBERT layer 1 shared substantial variance with spectrotemporal features and added little unique variance, but it did not capture all the articulatory or most of the semantic variance represented by the comparison features. HuBERT layer 9 explained variance beyond spectrotemporal features, shared substantial variance with articulations while the articulatory model added little or no unique variance, and shared more variance with word embeddings than layer 1 did. Word embeddings still explained unique variance in higher-level regions, while layer 1 better explained auditory cortex. These comparisons indicate that layer 9 combines brain-relevant low- and mid-level information with some lexical or semantic information, rather than representing only one processing level.

  5. Knowl 5 — Linear probes show linguistic information changing with model depth

    empirical result

    Probes of APC, wav2vec, wav2vec 2.0, and HuBERT showed that their layers differ in the acoustic and linguistic information they expose. In most models, lower layers best predicted FBANK spectral features, lower-middle layers best predicted spectrotemporal features, and upper-middle layers best represented word-level information. Phoneme identity was best captured by the uppermost layers of wav2vec 2.0 and HuBERT; the authors suggest this may reflect phoneme perception being influenced by word context. All layers in the four self-supervised models beat the most-frequent-category baseline on phoneme and word classification. For HuBERT, all layers also beat random-vector and shuffled-word-embedding baselines on the word-embedding probe. These results support the interpretation that speech models acquire representations at multiple linguistic levels across depth, despite lacking explicit phoneme or word representations.

  6. Knowl 6 — Voxel-wise ridge encoding models evaluate features on held-out speech

    model/method

    The study fitted a separate linear ridge-regression encoding model for each cortical voxel, mapping a stimulus feature sequence to that voxel's fMRI BOLD response. The ridge parameter was selected independently per voxel using 50 cross-validation iterations; each iteration sampled 40 training-data chunks, each totaling more than four minutes, to account for temporal autocorrelation. Feature sequences were downsampled to the fMRI rate of 0.5 Hz using Lanczos resampling, and each model's hemodynamic response was estimated with a four-delay finite-impulse-response model. For bidirectional neural networks, representations were extracted at the end of a 64-second sliding window with a 10-millisecond stride to enforce causality. Models were fitted on 26 stories (5.4 hours) and evaluated on one held-out 10-minute story; performance was the correlation between actual and predicted voxel responses.

  7. Knowl 7 — Speech representations and comparison feature spaces

    model/method

    The study extracted the encoder output and contextualizer layers from four self-supervised speech models pretrained on 960 hours of LibriSpeech without fine-tuning on annotated samples. APC takes log-Mel spectrograms and uses a three-layer GRU contextualizer with an autoregressive objective. Wav2vec takes waveforms and uses a seven-layer CNN encoder and twelve-layer CNN contextualizer with a contrastive objective. Wav2vec 2.0 BASE uses a seven-layer CNN encoder and twelve-layer Transformer contextualizer with masked contrastive training; HuBERT BASE uses the same encoder and contextualizer sizes with masked predictive training. A supervised Deep Speech 2 model, trained on the same 960 hours with a CTC objective, served as a neural-network comparison. Other feature spaces included Mel-filterbank (FBANK) and spectrotemporal acoustic features, hand-labeled phonemes mapped to 14 articulatory features, 985-dimensional word embeddings, and contextual features from GPT's ninth layer.

  8. Knowl 8 — Pairwise variance partitioning separates shared and unique predictions

    model/method

    To compare two feature spaces, the authors fitted an encoding model for each space separately and a joint model using their concatenated features. For voxel vv, let ρA,v\rho_{A,v} and ρB,v\rho_{B,v} be the prediction correlations from the separate models, and let ρA∪B,v\rho_{A\cup B,v} be the correlation from the joint model. Explained response variance was approximated by signed squared correlation, Q(ρ)=ρ2sgn⁡(ρ)Q(\rho)=\rho^2\operatorname{sgn}(\rho). The shared component was calculated as QA∩B,v=Q(ρA,v)+Q(ρB,v)−Q(ρA∪B,v)Q_{A\cap B,v}=Q(\rho_{A,v})+Q(\rho_{B,v})-Q(\rho_{A\cup B,v}); the unique component for feature space AA was QA∖B,v=Q(ρA,v)−QA∩B,vQ_{A\setminus B,v}=Q(\rho_{A,v})-Q_{A\cap B,v}, with the analogous expression for BB. The corresponding partial correlations were the square roots of these variance components. The analyses were restricted to pairwise comparisons of HuBERT layer 1 or layer 9 with a baseline, using banded ridge regression so the feature spaces in a joint model could have different regularization parameters.

  9. Knowl 9 — Linguistic probes test what each speech-model layer represents

    experimental setup

    The authors aligned representations from each self-supervised model with FBANK features, spectrotemporal features, phoneme identity, and word identity using the 27 narrative stories. For lower-rate target features, model representations were mean-pooled over time. Ridge regression mapped model layers to FBANK and spectrotemporal features; linear classifiers were used for phoneme and word identity. Probe scores were feature-averaged correlation for acoustic targets, classification accuracy for phonemes, and perplexity for words. Word-level tasks used 100-dimensional GloVe embeddings and their vocabulary. Each probe used an 80/10/10 story-level train/validation/test split and was run with three random seeds. The authors also report that linear MLP probes with a bottleneck layer did not change some of the probe conclusions.

  10. Knowl 10 — Naturalistic fMRI speech dataset

    experimental setup

    The fMRI dataset comprised seven healthy, fluent English-speaking participants (three female) who passively listened to 27 spoken stories from The Moth Radio Hour. The stories contained approximately 57,900 words and more than five hours of speech. Whole-brain BOLD data were collected every two seconds (a 0.5-Hz sampling rate). Transcripts were aligned to the audio with a forced aligner to obtain word and phoneme timings. Twenty-six stories formed the encoding-model training set, and one approximately 10-minute story was held out for evaluation.

Coverage note — No substantial contributed material was omitted; supplemental MRI acquisition, preprocessing, and ROI-localization procedures were left out because they support the experiments rather than constitute a distinct finding.

References

  1. 1.Alain, G. and Bengio, Y. Understanding intermediate layers using linear classifier probes. February 2017.
  2. 2.Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., Chen, G., Chen, J., Chen, J., Chen, Z., Chrzanowski, M., Coates, A., Diamos, G., Ding, K., Du, N., Elsen, E., Engel, J., Fang, W., Fan, L., Fougner, C., Gao, L., Gong, C., Hannun, A., Han, T., Johannes, L. V., Jiang, B., Ju, C., Jun, B., LeGresley, P., Lin, L., Liu, J., Liu, Y., Li, W., Li, X., Ma, D., Narang, S., Ng, A., Ozair, S., Peng, Y., Prenger, R., Qian, S., Quan, Z., Raiman, J., Rao, V., Satheesh, S., Seetapun, D., Sengupta, S., Srinet, K., Sriram, A., Tang, H., Tang, L., Wang, C., Wang, J., Wang, K., Wang, Y., Wang, Z., Wang, Z., Wu, S., Wei, L., Xiao, B., Xie, W., Xie, Y., Yogatama, D., Yuan, B., Zhan, J., and Zhu, Z. Deep speech 2: End-to-end speech recognition in English and mandarin. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML’16, pp. 173–182, New York, NY, USA, June 2016. JMLR.org.
  3. 3.Antonello, R., Turek, J. S., Vo, V. A., and Huth, A. Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses. In Advances in Neural Information Processing Systems, May 2021.
  4. 4.Baevski, A., Zhou, H., Mohamed, A., and Auli, M. Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. arXiv:2006.11477 [cs, eess], September 2020.
  5. 5.Binder, J. R., Desai, R. H., Graves, W. W., and Conant, L. L. Where Is the Semantic System? A Critical Review and Meta-Analysis of 120 Functional Neuroimaging Studies. Cerebral Cortex, 19(12):2767–2796, December 2009. ISSN 1047-3211. doi: 10.1093/cercor/bhp055.
  6. 6.Caucheteux, C., Gramfort, A., and King, J.-R. GPT-2’s activations predict the degree of semantic comprehension in the human brain, September 2021.
  7. 7.Chi, T., Gao, Y., Guyton, M. C., Ru, P., and Shamma, S. Spectro-temporal modulation transfer functions and speech intelligibility. The Journal of the Acoustical Society of America, 106(5):2719, October 1999. ISSN 0001-4966. doi: 10.1121/1.428100.
  8. 8.Chi, T., Ru, P., and Shamma, S. A. Multiresolution spectrotemporal analysis of complex sounds. The Journal of the Acoustical Society of America, 118(2):887–906, August 2005. ISSN 0001-4966. doi: 10.1121/1.1945807.
  9. 9.Chung, Y.-A. and Glass, J. Generative Pre-Training for Speech with Autoregressive Predictive Coding. arXiv:1910.12607 [cs, eess], January 2020.
  10. 10.Chung, Y.-A., Hsu, W.-N., Tang, H., and Glass, J. An Unsupervised Autoregressive Model for Speech Representation Learning. arXiv:1904.03240 [cs, eess], June 2019.
  11. 11.Dale, A. M., Fischl, B., and Sereno, M. I. Cortical Surface-Based Analysis: I. Segmentation and Surface Reconstruction. NeuroImage, 9(2):179–194, February 1999. ISSN 1053-8119. doi: 10.1006/nimg.1998.0395.
  12. 12.de Heer, W. A., Huth, A. G., Griffiths, T. L., Gallant, J. L., and Theunissen, F. E. The Hierarchical Cortical Organization of Human Speech Processing. The Journal of Neuroscience, 37(27):6539–6557, July 2017. ISSN 0270-6474, 1529-2401. doi: 10.1523/JNEUROSCI.3267-16.2017.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423.
  14. 14.Elman, J. L. and McClelland, J. L. Cognitive penetration of the mechanisms of perception: Compensation for coarticulation of lexically restored phonemes. Journal of Memory and Language, 27(2):143–165, April 1988. ISSN 0749-596X. doi: 10.1016/0749-596X(88)90071-X.
  15. 15.Ettinger, A., Elgohary, A., and Resnik, P. Probing for semantic evidence of composition by means of simple classification tasks. In Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, pp. 134–139, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/W16-2524.
  16. 16.Gao, J. S., Huth, A. G., Lescroart, M. D., and Gallant, J. L. Pycortex: An interactive surface visualizer for fMRI. Frontiers in Neuroinformatics, 9, 2015. ISSN 1662-5196.
  17. 17.Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A., Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, S. C., Casto, C., Fanda, L., Doyle, W., Friedman, D., Dugan, P., Melloni, L., Reichart, R., Devore, S., Flinker, A., Hasenfratz, L., Levy, O., Hassidim, A., Brenner, M., Matias, Y., Norman, K. A., Devinsky, O., and Hasson, U. Thinking ahead: Spontaneous prediction in context as a keystone of language in humans and machines. Preprint, Neuroscience, December 2020.
  18. 18.Hewitt, J. and Liang, P. Designing and Interpreting Probes with Control Tasks. arXiv:1909.03368 [cs], September 2019.
  19. 19.Hickok, G. and Poeppel, D. The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5):393–402, May 2007. ISSN 1471-0048. doi: 10.1038/nrn2113.
  20. 20.Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. arXiv:2106.07447 [cs, eess], June 2021.
  21. 21.Huth, A. G., de Heer, W. A., Griffiths, T. L., Theunissen, F. E., and Gallant, J. L. Natural speech reveals the semantic maps that tile human cerebral cortex. Nature, 532(7600):453–458, April 2016. ISSN 1476-4687. doi: 10.1038/nature17637.
  22. 22.Jain, S. and Huth, A. Incorporating Context into Language Encoding Models for fMRI. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 31, pp. 6628–6637. Curran Associates, Inc., 2018.
  23. 23.Jain, S., Vo, V., Mahto, S., LeBel, A., Turek, J. S., and Huth, A. Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech. Advances in Neural Information Processing Systems, 33:13738–13749, 2020.
  24. 24.Kell, A. J. E., Yamins, D. L. K., Shook, E. N., Norman-Haignere, S. V., and McDermott, J. H. A Task-Optimized Neural Network Replicates Human Auditory Behavior, Predicts Brain Responses, and Reveals a Cortical Processing Hierarchy. Neuron, 98(3):630–644.e16, May 2018. ISSN 0896-6273. doi: 10.1016/j.neuron.2018.03.044.
  25. 25.LeBel, A., Jain, S., and Huth, A. G. Voxelwise Encoding Models Show That Cerebellar Language Representations Are Highly Conceptual. Journal of Neuroscience, 41(50):10341–10355, December 2021. ISSN 0270-6474, 1529-2401. doi: 10.1523/JNEUROSCI.0118-21.2021.
  26. 26.Lieberman, P. Towards a Unified Phonetic Theory. Linguistic Inquiry, 1(3):307–322, 1970. ISSN 0024-3892.
  27. 27.Mesgarani, N., Cheung, C., Johnson, K., and Chang, E. F. Phonetic Feature Encoding in Human Superior Temporal Gyrus. Science, 343(6174):1006–1010, February 2014. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.1245994.
  28. 28.Millet, J. and King, J.-R. Inductive biases, pretraining and fine-tuning jointly account for brain responses to speech. arXiv:2103.01032 [cs, eess, q-bio], February 2021.
  29. 29.Norman-Haignere, S. V. and McDermott, J. H. Neural responses to natural and model-matched stimuli reveal distinct computations in primary and nonprimary auditory cortex. PLOS Biology, 16(12):e2005127, December 2018. ISSN 1545-7885. doi: 10.1371/journal.pbio.2005127.
  30. 30.Nunez-Elizalde, A. O., Huth, A. G., and Gallant, J. L. Voxelwise encoding models with non-spherical multivariate normal priors. NeuroImage, 197:482–492, August 2019. ISSN 1053-8119. doi: 10.1016/j.neuroimage.2019.04.012.
  31. 31.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5206–5210, South Brisbane, Queensland, Australia, April 2015. IEEE. ISBN 978-1-4673-6997-8. doi: 10.1109/ICASSP.2015.7178964.
  32. 32.Pasad, A., Chou, J.-C., and Livescu, K. Layer-wise Analysis of a Self-supervised Speech Representation Model. arXiv:2107.04734 [cs, eess], October 2021.
  33. 33.Pennington, J., Socher, R., and Manning, C. Glove: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543, Doha, Qatar, 2014. Association for Computational Linguistics. doi: 10.3115/v1/D14-1162.
  34. 34.Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 2227–2237, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1202.
  35. 35.Pisoni, D. B. and Sawusch, J. R. Some Stages of Processing in Speech Perception. In Cohen, A. and Nooteboom, S. G. (eds.), Structure and Process in Speech Perception, Communication and Cybernetics, pp. 16–35, Berlin, Heidelberg, 1975. Springer. ISBN 978-3-642-81000-8. doi: 10.1007/978-3-642-81000-8 2.
  36. 36.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving Language Understanding by Generative Pre-Training. pp. 12, 2018.
  37. 37.Schneider, S., Baevski, A., Collobert, R., and Auli, M. Wav2vec: Unsupervised Pre-training for Speech Recognition. arXiv:1904.05862 [cs], September 2019.
  38. 38.Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences, 118(45), November 2021. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.2105646118.
  39. 39.Shi, X., Padhi, I., and Knight, K. Does String-Based Neural MT Learn Source Syntax? In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1526–1534, Austin, Texas, 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1159.
  40. 40.Toneva, M. and Wehbe, L. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). In Advances in Neural Information Processing Systems, pp. 14954–14964, 2019.
  41. 41.Venezia, J. H., Thurman, S. M., Richards, V. M., and Hickok, G. Hierarchy of speech-driven spectrotemporal receptive fields in human auditory cortex. NeuroImage, 186:647–666, February 2019. ISSN 1053-8119. doi: 10.1016/j.neuroimage.2018.11.049.
  42. 42.Wehbe, L., Vaswani, A., Knight, K., and Mitchell, T. Aligning context-based statistical models of language with brain activity during reading. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 233–243, Doha, Qatar, October 2014. Association for Computational Linguistics. doi: 10.3115/v1/D14-1030.
  43. 43.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6.
  44. 44.Wu, M. C.-K., David, S. V., and Gallant, J. L. Complete functional characterization of sensory neurons by system identification. Annual Review of Neuroscience, 29(1):477–505, July 2006. ISSN 0147-006X, 1545-4126. doi: 10.1146/annurev.neuro.29.051605.113024.
  45. 45.Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., Huang, T.-H., Tseng, W.-C., Lee, K.-t., Liu, D.-R., Huang, Z., Dong, S., Li, S.-W., Watanabe, S., Mohamed, A., and Lee, H.-y. SUPERB: Speech processing Universal PERformance Benchmark. arXiv:2105.01051 [cs, eess], October 2021.
  46. 46.Zhuang, C., Yan, S., Nayebi, A., Schrimpf, M., Frank, M. C., DiCarlo, J. J., and Yamins, D. L. K. Unsupervised neural network models of the ventral visual stream. Proceedings of the National Academy of Sciences, 118(3), January 2021. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.2014196118.

Citation

MLA
Vaidya, A. R., et al. “Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech”. International Conference on Machine Learning, vol. 162, 2022, pp. 21927–44, https://proceedings.mlr.press/v162/vaidya22a.html.
APA
Vaidya, A. R., Jain, S., & Huth, A. (2022). Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech. International Conference on Machine Learning, 162, 21927–21944. https://proceedings.mlr.press/v162/vaidya22a.html
Chicago
Vaidya, A. R., S. Jain, and A. Huth. 2022. “Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech”. International Conference on Machine Learning 162: 21927–44. https://proceedings.mlr.press/v162/vaidya22a.html.
Harvard
Vaidya, A.R., Jain, S. and Huth, A. (2022) “Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech”, International Conference on Machine Learning. PMLR, pp. 21927–21944. Available at: https://proceedings.mlr.press/v162/vaidya22a.html.
Vancouver
1. Vaidya AR, Jain S, Huth A (2022) Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech. In: International Conference on Machine Learning. PMLR, pp 21927–21944

BibTeX

@InProceedings{pmlr-v162-vaidya22a,
  title = 	 {Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech},
  author =       {Vaidya, Aditya R and Jain, Shailee and Huth, Alexander},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {21927--21944},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/vaidya22a/vaidya22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/vaidya22a.html},
  abstract = 	 {Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models’ ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/