Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?

Subba Reddy OotaJashn AroraVeeral AgarwalMounika MarreddyManish GuptaBapi Raju Surampudi

article2022NAACL58 citations

Reveals how fine-tuning Transformers on specific NLP tasks improves fMRI brain response predictions across reading and listening modalities, identifying which syntactic and semantic objectives best match cortical activity in different brain regions.

Listen

Understanding how the human brain processes language is critical for advancing cognitive neuroscience and designing biologically inspired artificial intelligence. While recent natural language processing algorithms based on Transformers can predict brain activity, they are typically applied as generic, task-agnostic systems. Consequently, researchers have lacked insight into which specific language processing tasks—such as tracking sentence structure or comprehending high-level narrative meaning—most accurately correspond to neural activations during natural communication.

The article evaluates which specialized natural language processing task representations best predict functional magnetic resonance imaging brain responses. It systematically compares brain encoding models across two distinct modes of communication: reading written passages and listening to spoken stories.

To conduct this evaluation, the authors used fine-tuned Transformer models representing ten language tasks spanning syntactic and semantic functions, including coreference resolution, named entity recognition, shallow syntactic parsing, summarization, paraphrase detection, and natural language inference. Latent features from each task model were extracted and mapped via ridge regression to voxel-level brain data from two benchmark neuroimaging sources: the Pereira reading dataset (5 subjects, 627 sentences across 9 brain regions) and the Narratives listening dataset (primarily 82 subjects listening to naturalistic audio stories across 10 auditory and language regions). The models were evaluated using 10-fold cross-validation, measuring prediction accuracy, Pearson correlation, and mean absolute error.

The investigation produced four central findings. First, task specialization significantly influences brain encoding: models tuned for specific tasks consistently outperformed general, task-agnostic baseline models across both datasets. Second, the type of language task that best predicts brain activity depends heavily on the communication mode; shallow syntactic and entity-tracking tasks (coreference resolution, named entity recognition, and shallow syntax) yielded the highest predictive power for reading, whereas complex semantic reasoning tasks (paraphrase detection, summarization, and natural language inference) proved most predictive for spoken narrative comprehension. Third, predicted brain responses aligned with established biological hierarchies; reading showed strong engagement in the left language hemisphere and visual object areas, while listening engaged higher-order cognitive zones (such as the posterior medial cortex, achieving correlation coefficients up to approximately 0.23) far more than early auditory areas. Fourth, hierarchical clustering revealed that mathematical similarity structures among natural language processing models closely mirror the clustering of neural activation patterns observed across human subjects.

These findings indicate that human comprehension dynamically recruits distinct cognitive operations depending on whether language is read or heard. Listening to extended speech requires continuous semantic synthesis and abstraction, matching high-level natural language tasks, whereas reading relies more heavily on local syntactic and entity-binding structures. For decision-makers and technology developers, these results provide a concrete roadmap for selecting task-specific architectures when engineering brain-computer interfaces, neural prosthetics, and cognitively aligned artificial intelligence models.

Moving forward, researchers and developers should prioritize domain-aligned representations rather than generic language models when designing neuro-computational systems. Next research steps should focus on testing more sophisticated, non-linear predictive architectures, expanding the range of investigated language tasks, and directly comparing reading and listening modalities using identical textual stimuli within the same experimental cohorts.

Readers should interpret these outcomes within the scope of the study's boundaries. The analysis relied on linear ridge regression and utilized pre-existing models trained on datasets of differing sizes, which introduces potential scaling biases. Additionally, the reading and listening datasets involved different subjects, parcellation atlases, and stimulus materials. Despite these constraints, the statistical significance of the results (p < 0.01 across brain regions) supports high confidence in the fundamental finding that language task specialization systematically reflects neural processing.

arXiv: 2205.01404

No sufficiently relevant recommendations were found.

Cover for Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?

Abstract

Several popular Transformer based language models have been found to be successful for text-driven brain encoding. However, existing literature leverages only pretrained text Transformer models and has not explored the efficacy of task-specific learned Transformer representations. In this work, we explore transfer learning from representations learned for ten popular natural language processing tasks (two syntactic and eight semantic) for predicting brain responses from two diverse datasets: Pereira (subjects reading sentences from paragraphs) and Narratives (subjects listening to the spoken stories). Encoding models based on task features are used to predict activity in different regions across the whole brain. Features from coreference resolution, NER, and shallow syntax parsing explain greater variance for the reading activity. On the other hand, for the listening activity, tasks such as paraphrase generation, summarization, and natural language inference show better encoding performance. Experiments across all 10 task representations provide the following cognitive insights: (i) language left hemisphere has higher predictive brain activity versus language right hemisphere, (ii) posterior medial cortex, temporo-parieto-occipital junction, dorsal frontal lobe have higher correlation versus early auditory and auditory association cortex, (iii) syntactic and semantic tasks display a good predictive performance across brain regions for reading and listening stimuli resp.

Citation

MLA
Oota, S. R., et al. “Neural Language Taskonomy: Which NLP Tasks Are the Most Predictive of fMRI Brain Activity?”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3220–37, https://doi.org/10.18653/v1/2022.naacl-main.235.
APA
Oota, S. R., Arora, J., Agarwal, V., Marreddy, M., Gupta, M., & Surampudi, B. (2022). Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3220–3237. https://doi.org/10.18653/v1/2022.naacl-main.235
Chicago
Oota, S. R., J. Arora, V. Agarwal, M. Marreddy, M. Gupta, and B. Surampudi. 2022. “Neural Language Taskonomy: Which NLP Tasks Are the Most Predictive of fMRI Brain Activity?”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3220–37. https://doi.org/10.18653/v1/2022.naacl-main.235.
Harvard
Oota, S.R. et al. (2022) “Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3220–3237. Available at: https://doi.org/10.18653/v1/2022.naacl-main.235.
Vancouver
1. Oota SR, Arora J, Agarwal V, Marreddy M, Gupta M, Surampudi B (2022) Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3220–3237

BibTeX

@inproceedings{oota-etal-2022-neural,
    title = "Neural Language Taskonomy: Which {NLP} Tasks are the most Predictive of f{MRI} Brain Activity?",
    author = "Oota, Subba Reddy  and
      Arora, Jashn  and
      Agarwal, Veeral  and
      Marreddy, Mounika  and
      Gupta, Manish  and
      Surampudi, Bapi",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.235/",
    doi = "10.18653/v1/2022.naacl-main.235",
    pages = "3220--3237"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/