Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?
Subba Reddy OotaJashn AroraVeeral AgarwalMounika MarreddyManish GuptaBapi Raju Surampudi
Reveals how fine-tuning Transformers on specific NLP tasks improves fMRI brain response predictions across reading and listening modalities, identifying which syntactic and semantic objectives best match cortical activity in different brain regions.
Understanding how the human brain processes language is critical for advancing cognitive neuroscience and designing biologically inspired artificial intelligence. While recent natural language processing algorithms based on Transformers can predict brain activity, they are typically applied as generic, task-agnostic systems. Consequently, researchers have lacked insight into which specific language processing tasks—such as tracking sentence structure or comprehending high-level narrative meaning—most accurately correspond to neural activations during natural communication.
The article evaluates which specialized natural language processing task representations best predict functional magnetic resonance imaging brain responses. It systematically compares brain encoding models across two distinct modes of communication: reading written passages and listening to spoken stories.
To conduct this evaluation, the authors used fine-tuned Transformer models representing ten language tasks spanning syntactic and semantic functions, including coreference resolution, named entity recognition, shallow syntactic parsing, summarization, paraphrase detection, and natural language inference. Latent features from each task model were extracted and mapped via ridge regression to voxel-level brain data from two benchmark neuroimaging sources: the Pereira reading dataset (5 subjects, 627 sentences across 9 brain regions) and the Narratives listening dataset (primarily 82 subjects listening to naturalistic audio stories across 10 auditory and language regions). The models were evaluated using 10-fold cross-validation, measuring prediction accuracy, Pearson correlation, and mean absolute error.
The investigation produced four central findings. First, task specialization significantly influences brain encoding: models tuned for specific tasks consistently outperformed general, task-agnostic baseline models across both datasets. Second, the type of language task that best predicts brain activity depends heavily on the communication mode; shallow syntactic and entity-tracking tasks (coreference resolution, named entity recognition, and shallow syntax) yielded the highest predictive power for reading, whereas complex semantic reasoning tasks (paraphrase detection, summarization, and natural language inference) proved most predictive for spoken narrative comprehension. Third, predicted brain responses aligned with established biological hierarchies; reading showed strong engagement in the left language hemisphere and visual object areas, while listening engaged higher-order cognitive zones (such as the posterior medial cortex, achieving correlation coefficients up to approximately 0.23) far more than early auditory areas. Fourth, hierarchical clustering revealed that mathematical similarity structures among natural language processing models closely mirror the clustering of neural activation patterns observed across human subjects.
These findings indicate that human comprehension dynamically recruits distinct cognitive operations depending on whether language is read or heard. Listening to extended speech requires continuous semantic synthesis and abstraction, matching high-level natural language tasks, whereas reading relies more heavily on local syntactic and entity-binding structures. For decision-makers and technology developers, these results provide a concrete roadmap for selecting task-specific architectures when engineering brain-computer interfaces, neural prosthetics, and cognitively aligned artificial intelligence models.
Moving forward, researchers and developers should prioritize domain-aligned representations rather than generic language models when designing neuro-computational systems. Next research steps should focus on testing more sophisticated, non-linear predictive architectures, expanding the range of investigated language tasks, and directly comparing reading and listening modalities using identical textual stimuli within the same experimental cohorts.
Readers should interpret these outcomes within the scope of the study's boundaries. The analysis relied on linear ridge regression and utilized pre-existing models trained on datasets of differing sizes, which introduces potential scaling biases. Additionally, the reading and listening datasets involved different subjects, parcellation atlases, and stimulus materials. Despite these constraints, the statistical significance of the results (p < 0.01 across brain regions) supports high confidence in the fundamental finding that language task specialization systematically reflects neural processing.
- Paper: Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech, Aditya R. Vaidya et al. (2022). Its self-supervised speech encoding models establish how neural representations can predict fMRI responses to natural stories, providing a direct methodological precursor to the source’s task-specific comparisons.
- Paper: Toward a realistic model of speech processing in the brain with self-supervised learning, Juliette Millet et al. (2022). Its comparison of self-supervised speech representations with brain responses supplies a closely related encoding-model framework for understanding the source’s evaluation of language-task representations.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Its analysis of how BERT layers encode syntax, entities, and coreference clarifies the linguistic task representations that the source tests against neural activity.
No sufficiently relevant recommendations were found.
