Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning

Fan YinJesse VigPhilippe LabanShafiq JotyCaiming XiongChien-Sheng Wu

article2023ACL58 citations

Reveals that language models ignore most natural-language task definition content beyond output label specifications, and introduces structured prompting and meta-tuning strategies that significantly boost generalization on unseen tasks while cutting instruction length by over half.

Listen

Large language models are increasingly trained on natural language task definitions to generalize across unseen tasks without task-specific retraining. However, authoring comprehensive, free-form instructions is labor-intensive, and it remains unclear whether models interpret these instructions as intended or if the added text simply introduces noise.

The article evaluates which components of natural language task definitions are necessary for model performance and tests whether structured, compressed formats can improve instruction-following efficiency.

The authors conducted an empirical evaluation using the English portion of the Super-NaturalInstructions benchmark across 876 diverse language tasks. They categorized definition texts into eight functional types to perform systematic ablation experiments on BART and T5 language models. In parallel, they evaluated an automated Syntax-guided Task Definition Compression algorithm to remove non-essential phrasing. Building on these analyses, they introduced two interventions: structuring task definitions into standardized key-value triplets (input, action, and output) and adding an intermediate meta-tuning training phase to help models map examples directly to these structured components.

The analysis produced three primary findings. First, detailed descriptions of task inputs and secondary constraints contribute little to generalization performance; removing them caused minimal degradation, while larger models showed only a modest ability to leverage these extra details. Second, output and label specifications are critical for classification tasks: providing label lists enables models to identify valid outputs for unseen tasks, while label descriptions allow models to disambiguate familiar terms used in new contexts. Third, automated compression demonstrated that approximately 60% of words in task definitions can be pruned without harming accuracy, actually boosting performance on held-out tasks by up to 2.8 ROUGE-L points. Finally, combining structured triplet definitions with meta-tuning generated consistent performance gains across 119 unseen test tasks, yielding improvements of 4.2 points on BART-Large and 2.1 points on T5-XL.

These results indicate that current instruction-tuning models rely heavily on specific core signals—primarily target outputs and labels—rather than processing full, natural language descriptions holistically. Replacing lengthy, unstructured text with concise metadata or standardized templates reduces developer effort, saves context-window budget, and enhances overall model performance across both small and large architectures.

Organizations developing or deploying instruction-tuned systems should transition from crafting lengthy, unstructured prose instructions to utilizing concise, structured templates that explicitly specify expected outputs and label sets. Teams should also adopt intermediate meta-tuning stages to prime models to recognize consistent structural formats before fine-tuning on downstream tasks.

Confidence in these findings is high for standard text classification and structured natural language processing tasks within English-language benchmarks. However, caution is advised when extending these conclusions to open-ended generation, multilingual workflows, or ultra-large parameter models, where emergent reasoning capabilities may utilize unstructured context differently.

Cover for Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning

Abstract

Large language models (LLMs) have shown impressive performance in following natural language instructions to solve unseen tasks. However, it remains unclear whether models truly understand task definitions and whether the human-written definitions are optimal. In this paper, we systematically study the role of task definitions in instruction learning. We first conduct an ablation analysis informed by human annotations to understand which parts of a task definition are most important, and find that model performance only drops substantially when removing contents describing the task output, in particular label information. Next, we propose an automatic algorithm to compress task definitions to a minimal supporting set of tokens, and find that 60% of tokens can be removed while maintaining or even improving model performance. Based on these results, we propose two strategies to help models better leverage task instructions: (1) providing only key information for tasks in a common structured format, and (2) adding a meta-tuning stage to help the model better understand the definitions. With these two strategies, we achieve a 4.2 Rouge-L improvement over 119 unseen test tasks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Ablation Analysis of Annotated Task Definitions
  • 3.1 Ablation Analysis
  • 3.2 The Role of Label Information
  • 4 Compressing Task Definitions
  • 5 Improving Model Understanding of Task Definitions
  • 5.1 Experiments
  • 6 Related Work
  • 7 Discussion
  • 8 Conclusion
  • 9 Limitations
  • Acknowledgements
  • References
  • A Dataset and Model Details
  • A.1 Validation Task set
  • A.2 Model Training
  • B Annotation Procedure Details
  • B.1 Overview of the Annotation Procedure
  • B.2 A Hierarchy of Content Types in Definitions
  • B.3 Inter-Annotator Agreement Level
  • B.4 Pre-process and Post-process of The Annotations
  • C Compression Algorithm
  • D Examples of Triplet
  • E Hyperparameters Tuning for Instruction Learning after Meta-Tuning

Knowls

  1. Knowl 1 — Task definitions were annotated with a hierarchy of eight information types

    definition

    The study annotated task-definition text in the English portion of NIv2 using eight categories: Input Content (the primary description of the input), Additional Input Content (extra input details or constraints), Input Mention (an input reference embedded in an action description), Action Content (the operation to perform), Output Content (the primary description of the output), Additional Output Content (extra output details or constraints), Label List (classification label verbalizers), and Label Definition (natural-language explanations of labels). The annotation schematic on page 3 illustrates how these categories apply to task-definition text. Three authors annotated all 876 task definitions; a shared set of 150 definitions was used to assess agreement, with overall Fleiss’ kappa reported as 0.91. Category-level kappas were 0.92 for Input Content, 0.98 for Action Content, 0.83 for Output Content, 0.88 for Label List, 0.84 for Label Definition, 0.87 for Additional Input Content, 0.94 for Additional Output Content, and 1.0 for Input Mention.

  2. Knowl 2 — Ablations identify label information as the strongest contributor to validation performance

    data/table

    The authors retrained BART-Large (400M), T5-Large (770M), and T5-XL (3B) with selected content types removed from task definitions, then evaluated Rouge-L on a held-out validation set of 76 NIv2 training tasks. Scores are means over three random seeds; each model is reported on all tasks, classification tasks, and generation tasks. The page-4 results table shows that removing label information particularly harms classification, whereas removing input descriptions or secondary details generally has little effect. Metadata-only instructions are competitive with full definitions. “% kept” is the fraction of definition tokens remaining after the stated removal.

    BART-Large T5-Large T5-XL
    Method % kept All Cls. Gen. All Cls. Gen. All Cls. Gen.
    Heuristics – 39.22 53.36 28.94 39.22 53.36 28.94 39.22 53.36 28.94
    No definition 0% 38.63 45.77 33.43 43.56 53.52 36.45 44.26 55.64 35.99
    Shuffled definition 100% 39.73 49.08 32.94 45.25 57.17 36.59 48.57 64.10 37.26
    Metadata – 40.48 52.70 31.58 46.79 59.27 37.71 53.21 73.43 39.24
    Full definition 100% 40.17 48.92 33.79 47.55 60.20 38.34 53.63 70.82 41.17
    Remove additional input 87% 40.07 48.84 33.68 48.58 61.28 39.26 51.96 67.00 40.03
    Remove additional output 69% 39.72 47.62 33.65 48.38 63.31 37.51 51.29 66.32 39.36
    Remove all additional details 56% 39.81 47.90 33.71 48.04 62.01 37.89 52.16 66.70 40.60
    Remove label list 92% 36.70 44.23 31.22 44.95 58.29 35.26 46.34 60.45 36.09
    Remove label definitions 89% 38.04 47.06 32.10 46.86 57.42 37.46 47.25 61.28 37.04
    Remove all label information 80% 36.99 42.79 32.78 43.58 55.14 35.17 43.85 55.30 35.53
    Remove all output information 34% 37.18 43.43 32.63 43.60 55.24 35.14 43.98 55.99 35.23
    Remove all input information 67% 39.75 48.85 33.14 50.01 64.69 39.33 51.61 64.94 41.92
  3. Knowl 3 — Label verbalizers and label definitions support different kinds of generalization

    empirical result

    For classification tasks, the authors split validation tasks according to whether their combined label verbalizers had appeared in training. The reported values are Rouge-L changes when the specified label information is removed; negative values indicate a performance loss. Removing the Label List had nearly no effect for tasks with seen verbalizers (+0.12), but reduced performance by 15.85 points for tasks with unseen verbalizers. Removing Label Definitions reduced performance for both groups: by 13.21 points for seen verbalizers and 6.09 points for unseen verbalizers. The pattern indicates that verbalizers help transfer to a new label space, while natural-language definitions clarify label meaning, including when familiar label names are used with task-specific semantics.

    Label verbalizers seen in training Remove Label List Remove Label Definition
    Seen 0.12 -13.21
    Unseen -15.85 -6.09
  4. Knowl 4 — Syntax-guided Task Definition Compression removes parse-tree constituents that do not hurt performance

    algorithm

    Syntax-guided Task Definition Compression (STDC) uses black-box model predictions and a constituency parse of a task definition to find a shorter definition that preserves task performance. For each task, the authors used 100 examples as the representative set for compression and a separate 100 examples from that task to assess generalization. The compression criterion is non-degradation relative to the model’s performance with the full definition on the representative examples.

    Input: Model f; representative examples D for one task; full definition X; constituency parse tree T.
    Output: Compressed definition Xcompressed.
    Measure baseline performance b = f's task score on D using X.
    Set Xcompressed = X.
    Traverse constituents of T from higher levels to lower levels.
    For each constituent node n in that order:
        If n has already been removed as part of a larger subtree, continue.
        Form candidate Xcandidate by deleting the text corresponding to n and its subtree from Xcompressed.
        Measure candidate performance p on D using Xcandidate.
        If p is at least b:
            Set Xcompressed = Xcandidate.
    Return Xcompressed.
  5. Knowl 5 — STDC shortened definitions while improving held-out performance across three models

    data/table

    STDC was applied at inference time to definitions for three models. The table reports the fraction of definition tokens retained, Rouge-L before and after compression on a separate 100-example holdout set per task, and the proportion of examples for which compression increased performance (“coverage”). All three models improved after substantial shortening; T5-XL retained 41% of definition tokens and gained 2.8 Rouge-L points. The roughly 90% coverage values show that gains were not confined to a small number of examples.

    Model Fraction retained Before Rouge-L After Rouge-L Coverage
    BART-Large 0.52 40.7 41.9 0.89
    T5-Large 0.34 47.7 49.3 0.92
    T5-XL 0.41 50.3 53.1 0.89
  6. Knowl 6 — Compressed definitions retain output and label information more often than input details

    empirical result

    An analysis of the T5-XL STDC outputs found sharply different retention rates across the human-annotated content categories. The page-6 chart reports the following fractions retained: Input Content 0.33, Additional Input Content 0.47, Input Mention 0.41, Action Content 0.52, Output Content 0.66, Additional Output Content 0.67, Label Definition 0.80, and Label List 0.43. This distribution aligns with the ablation results: output descriptions and especially label definitions are more likely to survive compression than input descriptions. The authors also observed that STDC often preserved only some label verbalizers, suggesting that models could infer the remainder of the label space from partial lists; action verbs, including root verbs, were frequently removable. Across task categories, generation tasks were generally compressed more aggressively than classification tasks.

  7. Knowl 7 — Structured task triplets express instructions as input, action, and output fields

    model/method

    The structured-definition strategy represents each task with a JSON-like triplet containing Task input, Task action, and Task output. The authors derived triplets from NIv2 task definitions using their human annotations and automated parsing. For the input field, they extracted a noun phrase from Input Content; for the action field, they extracted a verb phrase from Action Content. Classification outputs were represented with the task’s label verbalizers and label definitions. For generation tasks, the output noun was extracted from the action sentence with rule-based processing. The resulting triplets were manually inspected, with parsing errors and some coreference problems corrected. Thus, the method both standardizes the format and condenses the original definition to selected task information.

  8. Knowl 8 — Meta-tuning trains models to produce structured task fields from tags and examples

    model/method

    The meta-tuning stage adapts a model to the structured task-definition format before ordinary instruction learning. Each training instance contains one of three field tags—Task input, Task action, or Task output—together with two task demonstration examples; the target is the corresponding field from that task’s triplet. Each task therefore yields three meta-tuning instances. The same two demonstrations are used during meta-tuning and subsequent instruction learning, so the meta-tuning stage does not introduce extra examples. The model is trained on the training tasks with a maximum-likelihood objective, and its resulting parameters initialize instruction learning. In the reported setup, meta-tuning ran for 10 epochs with batch size 16 and a constant learning rate of 5×10−65\times10^{-6} for BART-Large or 1×10−51\times10^{-5} for T5 models.

  9. Knowl 9 — Triplet formatting and meta-tuning improve zero-shot performance on 119 unseen tasks

    data/table

    The official NIv2 test evaluation used 100 balanced examples for each of 119 unseen tasks and reports mean Rouge-L with standard deviation over three random seeds. The page-8 results table shows gains from triplet definitions for every tested model and additional gains from meta-tuning. Relative to the corresponding standard full-definition model, triplets plus meta-tuning improve BART-Large by 4.19 points, Tk-INSTRUCT-Large by 3.96 points, and Tk-INSTRUCT-XL by 2.04 points. The listed T0, InstructGPT, and heuristic scores provide benchmark context; the central comparisons are between each Tk-INSTRUCT model and its triplet-trained variants.

    Model and task-definition strategy Rouge-L
    Heuristics 38.61
    T0 (11B) 32.30
    InstructGPT (175B) 52.10
    BART-Large (full definition) 40.70±0.440.70\pm0.4
    BART-Large + triplet 43.76±0.343.76\pm0.3
    BART-Large + triplet + meta-tuning 44.89±0.344.89\pm0.3
    Tk-INSTRUCT-Large (770M) 47.50±0.247.50\pm0.2
    Tk-INSTRUCT-Large + triplet 50.84±0.150.84\pm0.1
    Tk-INSTRUCT-Large + triplet + meta-tuning 51.46±0.251.46\pm0.2
    Tk-INSTRUCT-XL (3B) 54.08±0.354.08\pm0.3
    Tk-INSTRUCT-XL + triplet 55.58±0.255.58\pm0.2
    Tk-INSTRUCT-XL + triplet + meta-tuning 56.12±0.256.12\pm0.2
  10. Knowl 10 — The study’s conclusions are limited to English task definitions and models up to 3B parameters

    limitation

    The analysis covers English-language NIv2 tasks and models no larger than 3B parameters; its findings may not transfer to other languages or larger models. The structured-triplet experiment reformats information extracted from existing definitions, rather than testing whether authors creating triplets from scratch would obtain the same benefit. The experiments focus mainly on typical NLP tasks, particularly classification, rather than open-ended user instructions. The authors also note that some effects remain unexplained, including the observation that removing classification label information during training can affect performance on generation tasks.

Coverage note — No substantial contributed result was omitted; illustrative per-task compression and triplet examples were left out because they demonstrate the methods but add no distinct generalizable finding.

References

  1. 1.Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  3. 3.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  4. 4.Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu. 2022. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  6. 6.Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018. Pathologies of neural models make interpretations difficult. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3719–3728.
  7. 7.Joseph L Fleiss, Bruce Levin, and Myunghee Cho Paik. 2013. Statistical methods for rates and proportions. john wiley & sons.
  8. 8.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830.
  9. 9.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms.
  10. 10.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689.
  11. 11.Po-Nien Kung and Nanyun Peng. 2023. Do models really learn to follow instructions? an empirical study of instruction tuning.
  12. 12.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  13. 13.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597.
  14. 14.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  15. 15.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114.
  16. 16.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  17. 17.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021. Reframing instructional prompts to gptk’s language. arXiv preprint arXiv:2109.07830.
  18. 18.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In ACL.
  19. 19.OpenAI. 2023. Chatgpt. https://openai.com/blog/chatgpt/. Accessed on May 3, 2023.
  20. 20.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  21. 21.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  22. 22.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2022. Grips: Gradient-free, edit-based instruction search for prompting large language models.
  23. 23.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  24. 24.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang A. Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Rose Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. ArXiv, abs/2110.08207.
  25. 25.Timo Schick and Hinrich Schutze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269.
  26. 26.Timo Schick and Hinrich Schutze. 2021b. Few-shot text generation with natural language instructions. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 390–402.
  27. 27.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235.
  28. 28.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  29. 29.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  30. 30.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022b. Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks. In EMNLP.
  31. 31.Albert Webson, Zhizhong Chen, Carsten Eickhoff, and Ellie Pavlick. 2020. Do “Undocumented Workers” == “Illegal Aliens”? Differentiating Denotation and Connotation in Vector Spaces. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4090–4105.
  32. 32.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  33. 33.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. ArXiv, abs/2109.01652.
  34. 34.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022b. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  35. 35.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45.
  36. 36.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. Crossfit: A few-shot learning challenge for cross-task generalization in nlp. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7163–7189.
  37. 37.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.

Citation

MLA
Yin, F., et al. “Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 3063–79, https://doi.org/10.18653/v1/2023.acl-long.172.
APA
Yin, F., Vig, J., Laban, P., Joty, S., Xiong, C., & Wu, C.-S. (2023). Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3063–3079. https://doi.org/10.18653/v1/2023.acl-long.172
Chicago
Yin, F., J. Vig, P. Laban, S. Joty, C. Xiong, and C.-S. Wu. 2023. “Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3063–79. https://doi.org/10.18653/v1/2023.acl-long.172.
Harvard
Yin, F. et al. (2023) “Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3063–3079. Available at: https://doi.org/10.18653/v1/2023.acl-long.172.
Vancouver
1. Yin F, Vig J, Laban P, Joty S, Xiong C, Wu C-S (2023) Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3063–3079

BibTeX

@inproceedings{yin-etal-2023-read,
    title = "Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction Learning",
    author = "Yin, Fan  and
      Vig, Jesse  and
      Laban, Philippe  and
      Joty, Shafiq  and
      Xiong, Caiming  and
      Wu, Chien-Sheng",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.172/",
    doi = "10.18653/v1/2023.acl-long.172",
    pages = "3063--3079"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/