Built independently by an author, for readers. Read the story and support ChapterPal

keyword

NLP tasks

NLP tasks are specific language-processing problems that a computer system is designed to perform on human language, such as classifying text, extracting information, translating, summarizing, answering questions, or generating and rewriting text. They may involve understanding or transforming language in written or spoken form.

4 items

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, Xudong Shen

OrganizationsAllen Institute for AIAmirkabir University of TechnologyArizona State UniversityColumbia UniversityFactored AIGovernment Polytechnic, RajkotIndian Institute of Technology KharagpurIndian Institute of Technology MadrasJohns Hopkins UniversityMicrosoftNational Institute of Technology KarnatakaNational University of SingaporePSG College of TechnologySharif University of TechnologyStanford UniversityTCS ResearchUniversity of AmsterdamUniversity of California BerkeleyUniversity of Massachusetts AmherstUniversity of WashingtonZycus Infotech

Why you should read this

Introduces Super-NaturalInstructions, a benchmark of over 1,600 diverse tasks with human-written instructions, and develops Tk-Instruct, a model that outperforms much larger baselines like InstructGPT on unseen task generalization.

How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, including but not limited to classification, extraction, infilling, sequence tagging, text rewriting, and text composition. This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions -- training models to follow instructions on a subset of tasks and evaluating them on the remaining unseen ones. Furthermore, we build Tk-Instruct, a transformer model trained to follow a variety of in-context instructions (plain language task definitions or k-shot examples). Our experiments show that Tk-Instruct outperforms existing instruction-following models such as InstructGPT by over 9% on our benchmark despite being an order of magnitude smaller. We further analyze generalization as a function of various scaling parameters, such as the number of observed tasks, the number of instances per task, and model sizes. We hope our dataset and model facilitate future progress towards more general-purpose NLP models.

Added

2026-10-04

Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?

Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?

Subba Reddy Oota, Jashn Arora, Veeral Agarwal, Mounika Marreddy, Manish Gupta, Bapi Raju Surampudi

OrganizationsINRIAInternational Institute of Information Technology, HyderabadMicrosoft

Why you should read this

Reveals how fine-tuning Transformers on specific NLP tasks improves fMRI brain response predictions across reading and listening modalities, identifying which syntactic and semantic objectives best match cortical activity in different brain regions.

Several popular Transformer based language models have been found to be successful for text-driven brain encoding. However, existing literature leverages only pretrained text Transformer models and has not explored the efficacy of task-specific learned Transformer representations. In this work, we explore transfer learning from representations learned for ten popular natural language processing tasks (two syntactic and eight semantic) for predicting brain responses from two diverse datasets: Pereira (subjects reading sentences from paragraphs) and Narratives (subjects listening to the spoken stories). Encoding models based on task features are used to predict activity in different regions across the whole brain. Features from coreference resolution, NER, and shallow syntax parsing explain greater variance for the reading activity. On the other hand, for the listening activity, tasks such as paraphrase generation, summarization, and natural language inference show better encoding performance. Experiments across all 10 task representations provide the following cognitive insights: (i) language left hemisphere has higher predictive brain activity versus language right hemisphere, (ii) posterior medial cortex, temporo-parieto-occipital junction, dorsal frontal lobe have higher correlation versus early auditory and auditory association cortex, (iii) syntactic and semantic tasks display a good predictive performance across brain regions for reading and listening stimuli resp.

Added

2026-10-03

The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions

The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions

Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han

OrganizationsMicrosoftUniversity of Illinois Urbana-Champaign

Why you should read this

Reveals a critical misalignment between academic NLP benchmarks and real-world needs by analyzing over 94,000 user-GPT interactions, identifying frequently requested yet neglected tasks such as planning, designing, and advising.

Recent progress in Large Language Models (LLMs) has produced models that exhibit remarkable performance across a variety of NLP tasks. However, it remains unclear whether the existing focus of NLP research accurately captures the genuine requirements of human users. This paper provides a comprehensive analysis of the divergence between current NLP research and the needs of real-world NLP applications via a large-scale collection of user-GPT conversations. We analyze a large-scale collection of real user queries to GPT. We compare these queries against existing NLP benchmark tasks and identify a significant gap between the tasks that users frequently request from LLMs and the tasks that are commonly studied in academic research. For example, we find that tasks such as “design” and “planning” are prevalent in user interactions but are largely neglected or different from traditional NLP benchmarks. We investigate these overlooked tasks, dissect the practical challenges they pose, and provide insights toward a roadmap to make LLMs better aligned with user needs.

Added

2026-10-03

Language Models are Unsupervised Multitask Learners

Language Models are Unsupervised Multitask Learners

Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever

OrganizationsOpenAI

Why you should read this

Demonstrates that a decoder-only Transformer scaled on diverse data acquires zero-shot capabilities without task-specific modifications.

Natural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on task-specific datasets. We demonstrate that language models begin to learn these tasks without any explicit supervision when trained on a new dataset of millions of webpages called WebText. When conditioned on a document plus questions, the answers generated by the language model reach 55 F1 on the CoQA dataset- matching or exceeding the performance of 3 out of 4 baseline systems without using the 127,000+ training examples. The capacity of the language model is essential to the success of zero-shot task transfer and increasing it improves performance in a log-linear fashion across tasks. Our largest model, GPT-2, is a 1.5B parameter Transformer that achieves state of the art results on 7 out of 8 tested language modeling datasets in a zero-shot setting but still underfits WebText. Samples from the model reflect these improvements and contain coherent paragraphs of text. These findings suggest a promising path towards building language processing systems which learn to perform tasks from their naturally occurring demonstrations.

Added

2026-08-04

License

Published with permission