The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Shayne LongpreLe HouTu VuAlbert WebsonHyung Won ChungYi TayDenny ZhouQuoc V. LeBarret ZophJason Wei

article2023ICML896 citations

Demonstrates how training with mixed zero-shot, few-shot, and chain-of-thought prompts drives superior instruction tuning performance, releasing the Flan collection to provide more effective and computationally efficient starting checkpoints for downstream tasks.

Listen

Large language models increasingly drive critical processing and automation tasks across organizations, yet their ability to follow diverse natural language instructions reliably and affordably remains a challenge. Instruction tuning—the process of training models on collections of formatted input-output tasks—has emerged as a key solution, but practitioners have lacked clear design guidelines on optimal dataset composition, prompt formats, and task balance. Consequently, existing models often require immense parameter scales or expensive, non-public human feedback pipelines to achieve high generalization performance.

The article systematically evaluates the core design decisions behind publicly available instruction-tuning methods and introduces the Flan 2022 Collection. By conducting controlled ablation studies on consistently sized models, the article demonstrates how specific data enrichment, task balancing, and prompt engineering strategies dramatically improve instruction-following capabilities across both familiar and completely unseen evaluation tasks.

To conduct this assessment, the authors trained a baseline 3-billion-parameter model using more than 1,800 diverse tasks consolidated from major public repositories, including Flan 2021, P3++, and Super-Natural Instructions. The evaluation isolated individual variables by stripping out specific techniques, such as task balance weighting, chain-of-thought reasoning tasks, input inversions, and prompt variations. Model performance was measured across standard held-in benchmarks, chain-of-thought reasoning tasks, and comprehensive held-out benchmarks, including the 57-subject Massive Multitask Language Understanding (MMLU) suite and BIG-Bench Hard (BBH).

The findings show that specific training techniques yield substantial performance gains. Combining zero-shot, few-shot, and chain-of-thought prompt templates during training provided a 2% or greater performance lift across all evaluation settings, with adding as little as 5% to 10% few-shot templates notably improving zero-shot accuracy. Enriched training techniques—such as inverting input-output pairs to generate questions from answers—alongside careful task-source balancing, produced 3% to 17% absolute accuracy improvements over competing open-source instruction-tuned models of equal size. Remarkably, the 3-billion-parameter Flan-T5 model outperformed much larger alternative systems, including the 175-billion-parameter OPT-IML-Max, on key held-out benchmarks. Additionally, task scaling experiments showed that while performance on familiar tasks peaked around 200 tasks, performance on unseen tasks continued to scale log-linearly up to 1,836 tasks.

These results demonstrate that smart data engineering and prompt structuring can substitute for raw parameter scale, significantly lowering the computing costs and infrastructure risks associated with deploying high-performing language models. Furthermore, when deployed as an initialization point for single downstream tasks, Flan-T5 converged much faster and achieved higher ultimate accuracy than standard pre-trained baselines. This positions instruction-tuned checkpoints as an environmentally and financially efficient standard starting point, reducing the recurring computational burden across enterprise fine-tuning pipelines.

Organizations developing or deploying language models should adopt instruction-tuned models like Flan-T5 as standard baseline checkpoints rather than starting from raw pre-trained models. Machine learning teams should implement mixed-prompt templates and input-inversion data augmentations in their internal fine-tuning workflows, while carefully balancing task sources instead of simply maximizing raw data volume.

Confidence in these findings is reinforced by rigorous, controlled ablations on a consistent 3-billion-parameter architecture. However, decision-makers should note that data source quality matters significantly; simply scaling task counts without maintaining diversity and balance can cause performance plateaus. Future work should explore more refined automated weighting mechanisms and assess how instruction generalization interacts with emerging synthetic data pipelines.

Cover for The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Abstract

We study the design decisions of publicly available instruction tuning methods, and break down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies on the Flan Collection of tasks and methods, we tease apart the effect of design decisions which enable Flan-T5 to outperform prior work by 3-17%+ across evaluation settings. We find task balancing and enrichment techniques are overlooked but critical to effective instruction tuning, and in particular, training with mixed prompt settings (zero-shot, few-shot, and chain-of-thought) actually yields stronger (2%+) performance in all settings. In further experiments, we show Flan-T5 requires less finetuning to converge higher and faster than T5 on single downstream tasks, motivating instruction-tuned models as more computationally-efficient starting checkpoints for new tasks. Finally, to accelerate research on instruction tuning, we make the Flan 2022 collection of datasets, templates, and methods publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Public Instruction Tuning Collections
  • 3 Flan 2022 Instruction Tuning Experiments
  • 3.1 Ablation Studies
  • 3.2 Training with Mixed Prompt Settings
  • 3.3 Scaling Small Models to 1.8k+ Tasks
  • 3.4 Task Enrichment with Input Inversion
  • 3.5 Balancing Data Sources
  • 3.6 Discussion
  • 4 Instruction Tuning Enhances Single-Task Finetuning
  • 5 Related Work
  • 6 Conclusions
  • References
  • A Experimental Details
  • A.1 Instruction Tuning
  • A.2 Single-Task Finetuning
  • A.3 Evaluation
  • B Input Inversion Details

Knowls

  1. Knowl 1 — Flan 2022 improves on earlier instruction-tuning collections

    empirical result

    In a controlled comparison using 3-billion-parameter T5-XL models, Flan 2022 generally scores higher than models tuned on Flan 2021, P3++, or Super-Natural Instructions. Each score pair below gives zero-shot / few-shot performance, in this order: Held-In accuracy, Chain-of-Thought (CoT) accuracy, MMLU accuracy, BIG-Bench Hard (BBH) accuracy, and BBH-CoT accuracy.

    Flan 2022 scores are 73.8 / 74.8, 35.8 / 34.1, 50.3 / 52.4, 26.2 / 39.3, and 33.9 / 35.2. Flan 2021 scores are 68.4 / 56.3, 24.6 / 22.7, 41.4 / 34.8, 28.1 / 28.3, and 26.0 / 26.9. P3++ scores are 70.5 / 62.8, 25.6 / 25.6, 46.1 / 34.1, 26.0 / 30.8, and 23.4 / 26.1. Super-Natural Instructions scores are 50.3 / 42.2, 13.8 / 14.3, 35.6 / 31.1, 10.4 / 15.6, and 8.0 / 12.5.

    Relative to the best alternative T5-XL collection in each setting, Flan 2022's score differences, in the same metric and prompt-setting order, are +3.3 / +12.0, +10.2 / +8.5, +4.2 / +17.6, −1.9 / +8.5, and +7.9 / +8.3. Thus, its advantage is especially large on few-shot MMLU and CoT tasks, while it does not lead on zero-shot BBH. The paper also reports Flan 2022 MMLU scores of 50.3 / 52.4 against OPT-IML-Max 175B's 49.1 / 47.1, and a few-shot BBH score of 39.3 against OPT-IML-Max 175B's 35.7; the latter comparison is not size-matched.

  2. Knowl 2 — Mixing prompt formats improves both zero-shot and few-shot performance

    empirical result

    Training a model on a mixture of zero-shot and few-shot prompt templates improves evaluation in both prompting modes, rather than forcing a trade-off between them. The experiments varied the fraction of few-shot templates used in training and evaluated accuracy on Held-In tasks and held-out MMLU.

    The reported curves show that adding as little as 5% few-shot templates can substantially improve zero-shot performance, while adding at least 10% zero-shot data improves few-shot performance. For both Held-In and MMLU evaluations, the strongest results occur across a broad range—roughly 10% to 90% few-shot training templates—and this mixed range consistently outperforms training with only one prompt setting. The paper also reports that adding 10% few-shot prompts can raise zero-shot performance by more than 2%.

  3. Knowl 3 — Ablations identify distinct contributions from four Flan training choices

    empirical result

    The authors separately removed CoT training data, input inversion, mixture balancing, or few-shot templates from Flan 2022 training. Each entry below is zero-shot / few-shot performance, with metrics ordered as Held-In, CoT, MMLU, BBH, and BBH-CoT. The full Flan 2022 results are 73.8 / 74.8, 35.8 / 34.1, 50.3 / 52.4, 26.2 / 39.3, and 33.9 / 35.2.

    • Without CoT training data: 73.3 / 73.2, 28.8 / 24.6, 47.5 / 46.9, 18.2 / 30.0, and 18.2 / 12.0.
    • Without input inversion: 73.8 / 74.1, 32.2 / 23.5, 41.7 / 41.2, 18.4 / 24.2, and 15.7 / 13.0.
    • Without mixture balancing: 71.2 / 73.1, 32.3 / 30.5, 45.4 / 45.8, 15.1 / 24.3, and 13.8 / 15.4.
    • Without few-shot templates: 72.5 / 62.2, 38.9 / 28.6, 47.3 / 38.7, 27.6 / 30.8, and 18.6 / 23.3.

    The pattern is specific to the method removed: CoT data is important for CoT evaluations; input inversion contributes strongly to held-out MMLU and BBH; few-shot templates especially support few-shot evaluation; and mixture balancing improves performance broadly.

  4. Knowl 4 — Task scaling benefits held-out performance, with different behavior by model size

    empirical result

    To test task scaling, the authors finetuned T5-LM-adapted models at Small, Base, Large, XL, and XXL sizes on randomly selected subsets of 8, 25, 50, 100, 200, 400, 800, or all available tasks. Every run included the Held-In tasks, allowing comparison of performance on tasks already represented during training. The paper describes the largest condition as 1,873 tasks in the subset list and reports results for 1,836 tasks in its discussion and figure.

    Held-In performance generally peaks around 200 tasks and then declines as more tasks are added; larger models peak later and lose less performance. Held-out MMLU performance rises approximately log-linearly with task count for the larger model sizes, with the best reported results at the all-task scale. Only T5-Small appears to exceed its held-out performance before the largest task condition. These results suggest that even T5-Base may not have exhausted its capacity at this scale, and that larger models might benefit from further task expansion.

  5. Knowl 5 — Task-source weighting matters beyond simply adding more tasks

    data/table

    The authors measured source contributions by removing one source at a time from an equally weighted task mixture. The values below are mean Held-In accuracy, mean CoT accuracy, and MMLU accuracy, respectively:

    • All sources, equal weighting: 64.9, 41.4, 47.3.
    • Without Flan 2021: 55.3, 38.6, 45.7.
    • Without T0-SF: 63.2, 43.4, 44.7.
    • Without Super-Natural Instructions: 65.9, 42.2, 46.8.
    • Without CoT tasks: 65.6, 29.1, 46.8.
    • Without program-synthesis tasks: 66.9, 42.3, 46.8.
    • Without dialog tasks: 65.4, 40.3, 47.1.
    • All sources, weighted mixture: 66.4, 40.1, 48.1.

    Removing Flan 2021 causes the largest drops in Held-In and MMLU scores among these omissions; removing T0-SF also lowers MMLU. Removing CoT tasks sharply reduces CoT accuracy, consistent with their specialized contribution. The weighted mixture improves Held-In and MMLU relative to the equal mixture, but its CoT score is lower. The authors used these source-ablation results to narrow the search for mixture weights, then selected weights using practitioner judgment.

  6. Knowl 6 — Input inversion augments task variety and helps held-out evaluations

    model/method

    Input inversion creates an additional supervised task by reversing an example's input and target: for a question-answering example, for instance, the model is given the answer and trained to produce the question. The Flan 2022 experiments added inverted examples for datasets not already represented with such tasks, including dialog, program synthesis, and CoT data. Dialog inversions prompt for conversation history from a current turn; program-synthesis inversions prompt for the coding question a program solves; and CoT inversions permute question, answer, and explanation so that one or more components become targets.

    The inverted examples were mixed with regular examples at a 30% rate—three inverted examples for every ten regular examples. In the ablations, removing input inversion left Held-In performance nearly unchanged but substantially reduced held-out MMLU and BBH results. The finding supports inversion as a useful task-enrichment method in this collection, particularly for generalization beyond the training tasks.

  7. Knowl 7 — Flan-T5 is a stronger and faster-converging starting point for single-task finetuning

    empirical result

    The authors compared direct target-task finetuning of T5-XL with target-task finetuning of Flan-T5 XL, and also measured Flan-T5 XL without further finetuning. On the seven plotted Held-In tasks, the reported score gains for Flan-T5 followed by target-task finetuning over direct T5 finetuning are ANLI +4.0, ARC +8.7, BoolQ +2.7, CosmosQA +1.0, RTE +7.8, SQuAD v2 +2.4, and AI2 Science +16.7. On five Held-Out tasks, the corresponding gains are CondaQA +2.6, CxC +0.0, MedNLI +0.1, PubmedQA +2.3, and WANLI +1.6. The paper characterizes the results across the tested Held-In and Held-Out tasks as a Pareto improvement over direct T5 finetuning; in some low-data cases, Flan-T5 without further target-task finetuning also exceeds the direct-T5 result.

    In convergence curves for WANLI, MedNLI, CondaQA, PubmedQA, and CxC, Flan-T5 reaches higher accuracy in fewer target-task training steps than T5. The single-task experiments used 100,000 training steps, constant learning rate 0.001, dropout probability 0.1, and batches of 128 sequences of length 512; checkpoints were saved every 20 steps, and the checkpoint with the best validation performance was evaluated. Tasks with fewer than 1,000 training examples were averaged over three random seeds.

  8. Knowl 8 — The Flan 2022 Collection combines public task sources with expanded templates and methods

    model/method

    Flan 2022 combines Flan 2021, P3++, and Super-Natural Instructions with additional reasoning, dialog, and program-synthesis datasets. The collection's timeline estimate is 1,836 tasks and 15 million examples; the paper cautions that task counts depend on how tasks are defined. The authors expanded template coverage and formatting variety, including variation in instruction placement, few-shot-example separators, and multiple-choice option formatting. During instruction tuning, templates span zero-shot, few-shot, and CoT prompting; few-shot and few-shot-CoT prompts use 2, 3, or 5 exemplars.

    The collection, templates, and data-generation methods are made publicly available. The released generation code allows researchers to vary source-mixture rates, templates, prompt types, and data-augmentation choices.

  9. Knowl 9 — Evaluation uses matched T5-XL models and separates seen, CoT, and held-out tasks

    experimental setup

    Unless otherwise stated, the instruction-tuning comparisons use the prefix-language-model-adapted T5-LM XL checkpoint, with 3 billion parameters, to keep model size consistent across collections. The Held-In score is mean accuracy over validation sets from four question-answering tasks and four natural-language-inference tasks: BoolQ, ARC Easy, ARC Challenge, AI2 Middle School Science, ANLI rounds 1–3, and RTE. CoT performance is mean accuracy across GSM8K, StrategyQA, SVAMP, Asdiv, and CommonsenseQA, evaluated with prompts requesting step-by-step explanations. Held-Out evaluation uses the 57 exams in MMLU and the 23 challenging tasks in BBH; MMLU tasks are excluded from Flan 2022 training to preserve the held-out status.

    For task-scaling experiments, the model-size comparison includes T5-LM-adapted Small, Base, Large, XL, and XXL models. Thus, the reported Held-In and Held-Out measures refer to different evaluation conditions: Held-In tasks are represented in the training collection, while MMLU and BBH are not.

  10. Knowl 10 — Attribution in comparisons with OPT-IML is not isolated

    limitation

    The paper's results do not identify a single cause for Flan 2022's advantage over OPT-IML. Although the authors estimate that about 94% of OPT-IML's tasks also appear in Flan 2022, and report broadly similar emphasis on major task sources, the comparison can still differ in pretraining, model architecture, task composition, example mixing, prompt-format mixtures, and template construction. The OPT-IML collection's full templates, processing, and example-mixing pipeline were not released, preventing a direct matched comparison. The authors consider template design and mixed prompt formats plausible major contributors, but note that their individual effects were not established by dedicated controlled experiments.

Coverage note — The detailed permutations of template formatting are included as collection design rather than separate knowls because the paper does not isolate their effects in dedicated controlled experiments.

References

  1. 1.Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. Muppet: Massive multi-task representations with pre-finetuning. In EMNLP, 2021. URL https://aclanthology.org/2021.emnlp-main.468.
  2. 2.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. arXiv e-prints, art. arXiv:2204.01691, April 2022.
  3. 3.Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q Tran, Dara Bahri, Jianmo Ni, et al. Ext5: Towards extreme multi-task scaling for transfer learning. arXiv preprint arXiv:2111.10952, 2021.
  4. 4.Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-jian Jiang, and Alexander Rush. PromptSource: An integrated development environment and repository for natural language prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 93–104, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-demo.9. URL https://aclanthology.org/2022.acl-demo.9.
  5. 5.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  6. 6.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  7. 7.Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. The fifth pascal recognizing textual entailment challenge. In TAC, 2009.
  8. 8.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. NeurIPS, 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, et al. PaLM: Scaling language modeling with Pathways. arXiv preprint arXiv:2204.02311, 2022. URL https://arxiv.org/abs/2204.02311.
  11. 11.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  12. 12.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  13. 13.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  14. 14.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  15. 15.Andrew M Dai and Quoc V Le. Semi-supervised sequence learning. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/file/7137debd45ae4d0ab9aa953017286b20-Paper.pdf.
  16. 16.Ashwin Devaraj, William Sheffield, Byron Wallace, and Junyi Jessy Li. Evaluating factuality in text simplification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7331–7345, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.506. URL https://aclanthology.org/2022.acl-long.506.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019. URL https://aclanthology.org/N19-1423.
  18. 18.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858, 2022.
  19. 19.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  20. 20.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. Attributed text generation via post-hoc research and revision. arXiv preprint arXiv:2210.08726, 2022.
  21. 21.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361, 2021.
  22. 22.Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  23. 23.Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P Bigham. Improving zero and few-shot generalization in dialogue through instruction tuning. arXiv preprint arXiv:2205.12673, 2022.
  24. 24.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=0RDcd5Axok.
  25. 25.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. ICLR, 2020. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  26. 26.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  27. 27.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689, 2022.
  28. 28.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. CoRR, abs/2106.09685, 2021. URL https://arxiv.org/abs/2106.09685.
  29. 29.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Cosmos qa: Machine reading comprehension with contextual commonsense reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2391–2401, 2019.
  30. 30.Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. Inner monologue: Embodied reasoning through planning with language models. In arXiv preprint arXiv:2207.05608, 2022.
  31. 31.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, Xian Li, Brian O’Horo, Gabriel Pereyra, Jeff Wang, Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, and Ves Stoyanov. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017, 2022. URL https://arxiv.org/abs/2212.12017.
  32. 32.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. PubMedQA: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2567–2577, 2019. URL https://aclanthology.org/D19-1259.
  33. 33.Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. Unifying question answering, text classification, and regression via span extraction. arXiv preprint arXiv:1904.09286, 2019.
  34. 34.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. UnifiedQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, 2020. URL https://aclanthology.org/2020.findings-emnlp.171.
  35. 35.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  36. 36.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  37. 37.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. EMNLP, 2021. doi: 10.18653/v1/2021.emnlp-main.243. URL https://aclanthology.org/2021.emnlp-main.243.
  38. 38.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.703. URL https://aclanthology.org/2020.acl-main.703.
  39. 39.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models, 2022. URL https://arxiv.org/abs/2206.14858.
  40. 40.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. Towards understanding and mitigating social biases in language models. In ICML, 2021.
  41. 41.Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. Wanli: Worker and ai collaboration for natural language inference dataset creation. arXiv preprint arXiv:2201.05955, 2022a. URL https://arxiv.org/abs/2201.05955.
  42. 42.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning, 2022b. URL https://arxiv.org/abs/2205.05638.
  43. 43.Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. Multi-task deep neural networks for natural language understanding. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4487–4496, 2019.
  44. 44.Shayne Longpre, Yu Wang, and Chris DuBois. How effective is task-agnostic data augmentation for pretrained transformers? In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4401–4411, 2020.
  45. 45.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063, 2021.
  46. 46.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.173. URL https://aclanthology.org/2020.acl-main.173.
  47. 47.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018.
  48. 48.Kris McGuffie and Alex Newhouse. The radicalization risks of gpt-3 and advanced neural language models. arXiv preprint arXiv:2009.06807, 2020.
  49. 49.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984, 2020.
  50. 50.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013. URL https://proceedings.neurips.cc/paper/2013/file/9aa42b31882ec039965f3c4923ce901b-Paper.pdf.
  51. 51.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. MetaICL: Learning to learn in context. In NAACL, 2022. URL https://aclanthology.org/2022.naacl-main.201.
  52. 52.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. Cross-task generalization via natural language crowdsourcing instructions. arXiv preprint arXiv:2104.08773, 2021.
  53. 53.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786, 2022.
  54. 54.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  55. 55.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial nli: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, 2020.
  56. 56.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022. URL https://arxiv.org/abs/2203.02155.
  57. 57.Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang. Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 2855–2870, 2021. URL https://aclanthology.org/2021.eacl-main.249.
  58. 58.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, 2021.
  59. 59.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. NAACL, 2018. URL https://aclanthology.org/N18-1202.
  60. 60.Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel Bowman. Intermediate-task transfer learning with pretrained language models: When and why does it work? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5231–5247, 2020.
  61. 61.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
  62. 62.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  63. 63.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020. URL https://arxiv.org/abs/1910.10683.
  64. 64.Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for squad. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, 2018.
  65. 65.Abhilasha Ravichander, Matt Gardner, and Ana Marasović. Condaqa: A contrastive reading comprehension dataset for reasoning about negation. arXiv preprint arXiv:2211.00295, 2022. URL https://arxiv.org/abs/2211.00295.
  66. 66.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189, 2022. URL https://arxiv.org/abs/2203.17189.
  67. 67.Alexey Romanov and Chaitanya Shivade. Lessons from natural language inference in the clinical domain. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1586–1596, 2018. URL https://aclanthology.org/D18-1187.
  68. 68.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. ICLR 2022, 2021. URL https://arxiv.org/abs/2110.08207.
  69. 69.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1339. URL https://aclanthology.org/D19-1339.
  70. 70.Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Aguera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan. Large language models encode clinical knowledge, 2022. URL https://arxiv.org/abs/2212.13138.
  71. 71.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022. URL https://arxiv.org/abs/2206.04615.
  72. 72.Mirac Suzgun, Nathan Scales, Nathaneal Scharli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny ZHou, and Jason Wei. Challenging BIG-Bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022. URL https://arxiv.org/abs/2210.09261.
  73. 73.Zeerak Talat, Aurélie Névéol, Stella Biderman, Miruna Clinciu, Manan Dey, Shayne Longpre, Alexandra Sasha Luccioni10, Maraim Masoud11, Margaret Mitchell10, Dragomir Radev12, et al. You reap what you sow: On the challenges of bias evaluation under multilingual settings. Challenges & Perspectives in Creating Large Language Models, page 26, 2022.
  74. 74.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, 2019.
  75. 75.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022a. URL https://arxiv.org/abs/2205.05131.
  76. 76.Yi Tay, Jason Wei, Hyung Won Chung, David R. So, Siamak Shakeri, Xavier Garcia, Vinh Q. Tran, Hauixiu Steven Zheng, Jinfeng Rao, Denny Zhou, Donald Metzler, Neil Houlsby, Quoc V. Le, and Mostafa Dehghani. Transcending scaling laws with 0.1% extra compute. In arxiv, 2022b.
  77. 77.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. LaMDA: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022. URL https://arxiv.org/abs/2201.08239.
  78. 78.Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer. Exploring and predicting transferability across NLP tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7882–7926, 2020. URL https://aclanthology.org/2020.emnlp-main.635.
  79. 79.Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer. SPoT: Better frozen model adaptation through soft prompt transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 5039–5059, 2022. URL https://aclanthology.org/2022.acl-long.346.
  80. 80.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2153–2162, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1221. URL https://aclanthology.org/D19-1221.
  81. 81.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  82. 82.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. What language model architecture and pretraining objective work best for zero-shot generalization? ICML, 2022a. URL https://arxiv.org/abs/2204.05832.
  83. 83.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instructions, 2022b. URL https://arxiv.org/abs/2212.10560.
  84. 84.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705, 2022c. URL https://arxiv.org/abs/2204.07705.
  85. 85.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. ICLR 2022, 2021. URL https://openreview.net/forum?id=gEZrGCozdqR.
  86. 86.Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813, 2022.
  87. 87.Zhiyang Xu, Ying Shen, and Lifu Huang. Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning, 2022. URL https://arxiv.org/abs/2212.10773.
  88. 88.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. Crossfit: A few-shot learning challenge for cross-task generalization in NLP. In EMNLP, 2021. URL https://arxiv.org/abs/2104.08835.
  89. 89.Seonghyeon Ye, Doyoung Kim, Joel Jang, Joongbo Shin, and Minjoon Seo. Guess the instruction! making language models stronger zero-shot learners. arXiv preprint arXiv:2210.02969, 2022.
  90. 90.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. Defending against neural fake news. Advances in neural information processing systems, 32, 2019.
  91. 91.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414, 2022.
  92. 92.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.

Citation

MLA
Longpre, S., et al. “The Flan Collection: Designing Data and Methods for Effective Instruction Tuning”. arXiv, 2023, http://arxiv.org/abs/2301.13688v2.
APA
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., & Roberts, A. (2023). The Flan Collection: Designing Data and Methods for Effective Instruction Tuning. arXiv. http://arxiv.org/abs/2301.13688v2
Chicago
Longpre, S., L. Hou, T. Vu, et al. 2023. “The Flan Collection: Designing Data and Methods for Effective Instruction Tuning”. arXiv. http://arxiv.org/abs/2301.13688v2.
Harvard
Longpre, S. et al. (2023) “The Flan Collection: Designing Data and Methods for Effective Instruction Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.13688v2.
Vancouver
1. Longpre S, Hou L, Vu T, et al (2023) The Flan Collection: Designing Data and Methods for Effective Instruction Tuning. arXiv

BibTeX

@article{longpre2023the,
  title = {The Flan Collection: Designing Data and Methods for Effective Instruction Tuning},
  author = {Longpre, Shayne and Hou, Le and Vu, Tu and Webson, Albert and Chung, Hyung Won and Tay, Yi and Zhou, Denny and Le, Quoc V. and Zoph, Barret and Wei, Jason and Roberts, Adam},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.13688v2},
  eprint = {2301.13688}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/