MetaICL: Learning to Learn In Context

Sewon MinMike LewisLuke ZettlemoyerHannaneh Hajishirzi

article2022NAACL601 citations

Proposes a meta-training framework that teaches language models how to perform in-context learning across diverse tasks, allowing smaller models to generalize to unseen target tasks and rival fully finetuned baselines without requiring task-specific templates or parameter updates.

Listen

Deploying language models to handle new natural language tasks usually requires either expensive, task-specific fine-tuning or zero-shot transfer methods that rely heavily on fragile, handcrafted prompt templates. Standard in-context learning offers a flexible alternative by letting models adapt simply through observing a few examples at test time without parameter updates. However, standard in-context performance regularly falls short of full fine-tuning and exhibits significant performance variance.

The article evaluates MetaICL, a meta-training framework designed to tune language models across a wide collection of diverse training tasks specifically to master in-context learning. The main objective is to demonstrate that explicitly training a model to infer task semantics from a few concatenated examples enables superior, reliable few-shot generalization on unseen tasks without relying on custom templates or parameter fine-tuning.

The researchers conducted an extensive empirical study across 142 datasets covering classification, question answering, natural language inference, and paraphrase detection, testing models on 52 unique target tasks across seven distinct evaluation splits. The experiments utilized a 770-million-parameter GPT-2 Large base model and benchmarked it against standard in-context learning, multi-task zero-shot baselines, task-specific fine-tuning, and a substantially larger 6-billion-parameter GPT-J model. Crucially, the approach eliminated manual prompt templates by using raw input-output pairs.

The primary finding is that MetaICL, particularly its noisy channel formulation, consistently outperforms existing in-context and zero-shot baselines across diverse benchmarks. In challenging shift scenarios such as moving from high-resource to low-resource tasks, natural language inference, and paraphrase detection, performance improved by 6 to 15 absolute percentage points over competitive baselines. Second, the 770-million-parameter MetaICL model matched or exceeded the performance of the 6-billion-parameter GPT-J baseline, effectively rivaling systems nearly eight times its size. Third, MetaICL approached and occasionally outperformed target models trained with supervised fine-tuning. Finally, ablation studies showed that dataset diversity during meta-training is critical to success, and while MetaICL excels without human-written instructions, combining MetaICL with natural language instructions yields even higher accuracy.

These findings indicate that language models can be taught general learning behaviors that transfer across disparate domains without requiring massive parameter counts or intensive manual prompt engineering. In operational contexts, this enables organizations to deploy smaller, more cost-effective models that adapt reliably to new tasks in real time without the computational overhead, storage burden, and deployment risks associated with maintaining separate fine-tuned models for every task.

Organizations aiming to build fast-adapting natural language processing pipelines should adopt multi-task meta-training over diverse high-quality task collections rather than investing heavily in custom prompt engineering. Where maximum accuracy is required, practitioners should combine meta-trained in-context learning with standardized task instructions. Future initiatives should focus on identifying optimal task combinations to maximize meta-training efficiency before full production deployment.

The primary limitations of this work are its restriction to classification and multiple-choice formats with fixed candidate sets, leaving open-ended text generation unaddressed. Additionally, in-context learning increases memory and compute demands at inference time due to longer input sequences containing concatenated examples. While confidence in the benchmarked task formats is high, further validation is necessary before applying this framework to generative tasks and large-scale enterprise deployments.

arXiv: 2110.15943
Cover for MetaICL: Learning to Learn In Context

Abstract

We introduce MetaICL (Meta-training for In-Context Learning), a new meta-training framework for few-shot learning where a pretrained language model is tuned to do in-context learning on a large set of training tasks. This meta-training enables the model to more effectively learn a new task in context at test time, by simply conditioning on a few training examples with no parameter updates or task-specific templates. We experiment on a large, diverse collection of tasks consisting of 142 NLP datasets including classification, question answering, natural language inference, paraphrase detection and more, across seven different meta-training/target splits. MetaICL outperforms a range of baselines including in-context learning without meta-training and multi-task learning followed by zero-shot transfer. We find that the gains are particularly significant for target tasks that have domain shifts from the meta-training tasks, and that using a diverse set of the meta-training tasks is key to improvements. We also show that MetaICL approaches (and sometimes beats) the performance of models fully finetuned on the target task, and outperforms much bigger models with nearly 8x parameters. Finally, we show that MetaICL is complementary to human-written instructions, and the best performance can be achieved by combining both approaches.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MetaICL
  • 3.1 Meta-training
  • 3.2 Inference
  • 3.3 Channel MetaICL
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Evaluation
  • 4.4 Experiment Details
  • 5 Experimental Results
  • 5.1 Main Results
  • 5.2 Ablations
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Dataset List
  • B Implementation Details
  • C Additional Results & Analyses
  • C.1 GPT-J results
  • C.2 Varying LM sizes
  • C.3 Which meta-training tasks are more helpful?
  • C.4 Does MetaICL generalize when semantic hints from label words are removed?
  • D Potential Risks

Knowls

  1. Knowl 1 — MetaICL: Meta-Training for In-Context Learning

    model/method

    MetaICL (Meta-training for In-Context Learning) is a meta-learning method where a pretrained language model is tuned on a diverse collection of source tasks to learn task conditioning directly from concatenated in-context demonstrations, without updating parameters at inference time.

    Let Ti={(xji,yji)}j=1Ni\mathcal{T}_i = \{(x_j^i, y_j^i)\}_{j=1}^{N_i} denote the training set of task ii from a set of CC meta-training tasks, where Ni≫kN_i \gg k. At each meta-training iteration:

    1. A task index i∈[1,C]i \in [1, C] is sampled uniformly.
    2. A sequence of k+1k+1 input-output pairs (x1,y1),…,(xk,yk),(xk+1,yk+1)(x_1, y_1), \dots, (x_k, y_k), (x_{k+1}, y_{k+1}) is drawn from Ti\mathcal{T}_i.
    3. The model receives the concatenation of the first kk labeled examples and the (k+1)(k+1)-th query input: x1,y1,…,xk,yk,xk+1x_1, y_1, \dots, x_k, y_k, x_{k+1}.
    4. The model parameters are updated by maximizing the log-likelihood of the (k+1)(k+1)-th target label: LMetaICL=log⁡P(yk+1∣x1,y1,…,xk,yk,xk+1)\mathcal{L}_{\text{MetaICL}} = \log P(y_{k+1} \mid x_1, y_1, \dots, x_k, y_k, x_{k+1})

    At test time, given an unseen target task with kk labeled training examples (x1,y1),…,(xk,yk)(x_1, y_1), \dots, (x_k, y_k), a test input xx, and a candidate output set C\mathcal{C} (e.g., classification labels or multiple-choice answer options), prediction is performed via: y^=arg⁡max⁡c∈CP(c∣x1,y1,…,xk,yk,x)\hat{y} = \arg\max_{c \in \mathcal{C}} P(c \mid x_1, y_1, \dots, x_k, y_k, x)

  2. Knowl 2 — Channel MetaICL

    model/method

    Channel MetaICL is a noisy channel formulation of MetaICL where the conditional probability P(y∣x)P(y \mid x) is reparameterized via Bayes' rule as P(y∣x)∝P(x∣y)P(y)P(y \mid x) \propto P(x \mid y) P(y). Assuming a uniform prior over the candidate label set C\mathcal{C} (P(y)=1∣C∣P(y) = \frac{1}{|\mathcal{C}|}), the model predicts the input text conditioned on the label.

    During meta-training, for k+1k+1 sampled examples (x1,y1),…,(xk+1,yk+1)(x_1, y_1), \dots, (x_{k+1}, y_{k+1}) from task ii, the language model is fed the concatenated sequence y1,x1,…,yk,xk,yk+1y_1, x_1, \dots, y_k, x_k, y_{k+1} and trained to generate xk+1x_{k+1} by maximizing: LChannel=log⁡P(xk+1∣y1,x1,…,yk,xk,yk+1)\mathcal{L}_{\text{Channel}} = \log P(x_{k+1} \mid y_1, x_1, \dots, y_k, x_k, y_{k+1})

    At inference time on an unseen target task, given kk labeled examples (x1,y1),…,(xk,yk)(x_1, y_1), \dots, (x_k, y_k), a test input xx, and candidate label set C\mathcal{C}, the predicted label is computed as: y^=arg⁡max⁡c∈CP(x∣y1,x1,…,yk,xk,c)\hat{y} = \arg\max_{c \in \mathcal{C}} P(x \mid y_1, x_1, \dots, y_k, x_k, c)

  3. Knowl 3 — Multi-Task Benchmark Setup and Meta-Training/Target Splits

    experimental setup

    The experimental evaluation spans 142 distinct English NLP datasets curated from CrossFit and UnifiedQA, evaluating zero-shot and few-shot cross-task transfer across 52 unique target tasks across seven disjoint meta-training/target splits:

    • High Resource →\to Low Resource (HR →\to LR): 61 meta-training tasks with ≥10,000\ge 10{,}000 training examples; 26 low-resource target tasks (<10,000< 10{,}000 examples).
    • Classification →\to Classification: 43 classification meta-training tasks; 20 classification target tasks.
    • Non-Classification →\to Classification: 37 non-classification meta-training tasks; 20 classification target tasks.
    • QA →\to QA: 37 question-answering meta-training tasks; 22 question-answering target tasks.
    • Non-QA →\to QA: 33 non-QA meta-training tasks; 22 question-answering target tasks.
    • Non-NLI →\to NLI: 55 non-NLI meta-training tasks; 8 NLI target tasks.
    • Non-Paraphrase →\to Paraphrase: 59 non-paraphrase meta-training tasks; 4 paraphrase target tasks.

    Across all splits, target tasks include subsets with complete domain shift (e.g., finance, poems, climate, medicine). MetaICL is template-free: inputs and labels use original dataset strings, separated by a single newline (\n), and demonstrations are separated by three newlines (\n\n\n). Target tasks are evaluated using k=16k=16 randomly sampled demonstrations over 5 random seeds, reporting Macro-F1 for classification and Accuracy for non-classification (macro-averaged average and worst-case across seeds).

  4. Knowl 4 — Cross-Task Few-Shot Performance Across Evaluation Splits

    data/table

    Using GPT-2 Large (770M parameters), MetaICL and Channel MetaICL outperform raw language model baselines, multi-task zero-shot baselines (k=0k=0), and match or exceed supervised fine-tuning on the target dataset without meta-training:

    Method HR→\toLR Class→\toClass non-Class→\toClass QA→\toQA non-QA→\toQA non-NLI→\toNLI non-Para→\toPara
    All Target Tasks
    0-shot 34.8 34.2 34.2 40.2 40.2 25.5 34.2
    Channel In-context 43.1 / 38.5 46.3 / 40.3 46.3 / 40.3 40.8 / 38.1 40.8 / 38.1 39.9 / 34.8 45.4 / 40.9
    Multi-task 0-shot 35.6 37.3 36.8 45.7 36.0 40.7 30.6
    Channel Multi-task 0-shot 38.8 40.9 42.2 42.1 36.4 36.8 35.1
    MetaICL 43.3 / 41.7 43.4 / 39.9 38.1 / 31.8 46.0 / 44.8 38.5 / 36.8 49.0 / 44.8 33.1 / 33.1
    Channel MetaICL 49.1 / 46.8 50.7 / 48.0 50.6 / 48.1 44.9 / 43.5 41.9 / 40.5 54.6 / 51.9 52.2 / 50.3
    Fine-tune 46.4 / 40.0 50.7 / 44.0 50.7 / 44.0 41.8 / 39.1 41.8 / 39.1 44.3 / 32.8 54.7 / 48.9
    Fine-tune w/ meta-train 52.0 / 47.9 53.5 / 48.5 51.2 / 44.9 46.7 / 44.5 41.8 / 39.5 57.0 / 44.6 53.7 / 46.9
    Target Tasks in Unseen Domains
    Channel In-context 39.6 / 33.6 39.6 / 33.6 39.6 / 33.6 44.7 / 40.6 44.7 / 40.6 40.4 / 35.7 44.1 / 36.8
    Multi-task 0-shot 35.4 28.0 28.6 71.2 40.3 33.5 35.0
    Channel MetaICL 47.7 / 44.7 41.9 / 37.8 48.0 / 45.2 57.9 / 56.6 47.2 / 45.0 62.0 / 57.3 51.0 / 49.9

    Entries report average / worst-case scores over 5 random seeds with k=16k=16. Channel MetaICL achieves the best average performance in 6 out of 7 settings, with large absolute improvements (6–15%) on low-resource and format-shifted splits (HR→\toLR, non-NLI→\toNLI, non-Para→\toPara).

  5. Knowl 5 — Parameter Efficiency and Comparison with 8x Larger Models

    empirical result

    MetaICL enables smaller models to match or exceed the few-shot performance of substantially larger raw language models:

    • Comparison to GPT-J (6B): GPT-2 Large (770M parameters) meta-trained with Channel MetaICL achieves higher or competitive performance compared to raw GPT-J (6B parameters, ~8x larger) evaluated with Channel In-context prompting across target benchmarks (e.g., HR→\toLR: 49.1% avg / 46.8% worst-case for GPT-2 Large Channel MetaICL vs. 48.6% avg / 44.4% worst-case for GPT-J Channel In-context; non-NLI→\toNLI: 54.6% vs. 47.2%; non-Para→\toPara: 52.2% vs. 51.0%).
    • Model Scaling: Evaluating GPT-2 sizes on HR→\toLR shows that Channel MetaICL on GPT-2 Small (124M parameters) achieves 46.2% average accuracy, outperforming raw GPT-2 XL (1.5B parameters, 12x larger) which achieves 43.5% with Channel In-context prompting.
  6. Knowl 6 — Complementarity of MetaICL and Natural Language Instructions

    empirical result

    In-context meta-training and human-written task instructions provide complementary benefits for cross-task generalization:

    Method w/o Instruct (0/task) w/ Instruct (1/task) w/ Instruct (8.3/task)
    0-shot 33.3 34.2 -
    In-context 34.5 / 31.5 45.2 / 42.3 -
    Channel In-context 39.0 / 35.4 39.6 / 35.3 -
    Multi-task 0-shot 35.7 32.6 37.1
    Channel Multi-task 0-shot 36.7 30.6 36.0
    MetaICL 40.4 / 37.7 42.6 / 41.0 43.2 / 41.0
    Channel MetaICL 42.2 / 40.0 45.3 / 43.9 46.9 / 44.2

    Key observations on 32 meta-training and 12 target tasks:

    1. MetaICL without instructions (40.4% direct / 42.2% channel) outperforms instruction-tuned multi-task 0-shot learning with 8.3 instructions per task (37.1% direct / 36.0% channel).
    2. Adding natural instructions to MetaICL yields additive gains, increasing Channel MetaICL performance from 42.2% (0 instructions) to 45.3% (1 instruction/task) and 46.9% (8.3 instructions/task).
    3. However, Channel MetaICL trained on all 61 available tasks without instructions reaches 49.1% (from the full HR→\toLR benchmark), outperforming instruction tuning restricted to the 32 tasks with human-authored prompts (46.9%).
  7. Knowl 7 — Impact of Meta-Training Task Diversity and Task Count

    empirical result

    The quantity and diversity of meta-training tasks strongly govern the downstream performance of MetaICL:

    • Task Quantity: Subsampling meta-training tasks in the HR→\toLR setting across C∈{7,15,30,61}C \in \{7, 15, 30, 61\} tasks demonstrates that downstream target performance steadily improves with larger task counts. However, variance across different random subsets of CC tasks is high.
    • Task Diversity: When holding the number of meta-training tasks fixed at 13:
      • A diverse set of 13 tasks (spanning QA, NLI, relation extraction, sentiment analysis, topic classification, and hate speech detection) achieves 45.6% / 43.4% (MetaICL) and 47.2% / 44.7% (Channel MetaICL).
      • A non-diverse set of 13 tasks (restricted to sentiment analysis, topic classification, and hate speech detection) achieves 38.8% / 35.4% (MetaICL) and 45.3% / 42.6% (Channel MetaICL).

    Task diversity is therefore a key determinant of cross-task transfer success.

  8. Knowl 8 — Demonstration Scaling and Context Truncation Dynamics

    empirical result

    Varying the number of in-context demonstrations k∈{0,4,8,16,32}k \in \{0, 4, 8, 16, 32\} during evaluation in the HR→\toLR setting reveals:

    1. Performance increases monotonically from k=0k=0 (zero-shot transfer) to k=16k=16 for both raw language models and meta-trained models.
    2. Channel MetaICL consistently outperforms raw in-context baselines across all demonstration shot counts kk.
    3. Performance gains saturate between k=16k=16 and k=32k=32. This saturation is driven by the maximum sequence length limit (1024 tokens), which necessitates truncating earlier demonstration examples when kk becomes large.
  9. Knowl 9 — In-Context Generalization Under Label Word Replacement

    empirical result

    When meaningful label tokens in classification tasks (e.g., entailment, not_entailment or positive, negative) are replaced by random English words sampled from a 61,569-word vocabulary:

    • Raw language models and models meta-trained on original label words drop to near-random accuracy (23.8%–37.7% vs. 36.0% random baseline), proving that standard few-shot prompting relies heavily on semantic priors in the label tokens.
    • Multi-task 0-shot models cannot handle replaced labels (achieving 23.1%–34.1%).
    • When MetaICL is meta-trained with independently randomized label mappings per iteration, it learns to bind label semantics dynamically from the kk context examples alone, improving performance on replaced labels from 30.1% to 43.5% (average) / 35.2% (worst-case), while retaining 40.7% accuracy on original label strings.
  10. Knowl 10 — Limitations of Meta-Trained In-Context Learning

    limitation

    MetaICL has three primary limitations:

    1. Inference and Training Efficiency: Feeding concatenated kk-shot demonstrations requires substantially longer context lengths (1024 tokens) than zero-shot transfer (256 tokens), incurring higher quadratic attention compute (O(L2)O(L^2)) and linear memory costs (O(L)O(L)) during training and inference.
    2. Task Scope: The formulation requires a predefined candidate output space C\mathcal{C} (e.g., categorical classification labels or multiple-choice answer candidates) for probability scoring and noisy-channel inversion, making it inapplicable to open-ended free-form text generation.
    3. Model Scale and Architecture Coverage: Meta-training experiments were conducted primarily on autoregressive causal language models of moderate size (GPT-2 Large, 770M parameters), without exploration on masked language models or 100B+ parameter models.

Coverage note — None was omitted; all key contributions, algorithms (MetaICL and Channel MetaICL), experimental configurations, main results, scaling analyses, label-randomization tests, instruction ablations, and stated limitations are represented.

References

  1. 1.Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021. Muppet: Massive multi-task representations with pre-finetuning. arXiv preprint arXiv:2101.11038.
  2. 2.Tiago A. Almeida, Jose Marıa G. Hidalgo, and Akebo Yamakami. 2011. Contributions to the study of sms spam filtering: New collection and results. In Proceedings of the 11th ACM Symposium on Document Engineering.
  3. 3.Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006. The second pascal recognising textual entailment challenge. In Proceedings of the second PASCAL challenges workshop on recognising textual entailment.
  4. 4.Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  5. 5.Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009. The fifth pascal recognizing textual entailment challenge. In TAC.
  6. 6.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In EMNLP.
  7. 7.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In ICLR.
  8. 8.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. In AAAI.
  9. 9.Michael Boratko, Xiang Li, Tim O'Gorman, Rajarshi Das, Dan Le, and Andrew McCallum. 2020. ProtoQA: A question answering dataset for prototypical common-sense reasoning. In EMNLP.
  10. 10.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In NeurIPS.
  11. 11.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21).
  12. 12.Ankush Chatterjee, Kedhar Nath Narahari, Meghana Joshi, and Puneet Agrawal. 2019. SemEval-2019 task 3: EmoContext contextual emotion detection in text. In Proceedings of the 13th International Workshop on Semantic Evaluation.
  13. 13.Michael Chen, Mike D'Arcy, Alisa Liu, Jared Fernandez, and Doug Downey. 2019. CODAH: An adversarially-authored question answering dataset for common sense. In Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP.
  14. 14.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact: A large-scale dataset for table-based fact verification. In ICLR.
  15. 15.Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, and He He. 2022. Meta-learning via language model in-context tuning. In ACL.
  16. 16.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In NAACL-HLT.
  17. 17.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  18. 18.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop.
  19. 19.Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, and Matt Gardner. 2019. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In EMNLP.
  20. 20.Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the 11th International AAAI Conference on Web and Social Media.
  21. 21.Ona de Gibert, Naiara Perez, Aitor Garcıa-Pablos, and Montse Cuadros. 2018. Hate Speech Dataset from a White Supremacy Forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2).
  22. 22.Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. Proceedings of Sinn und Bedeutung.
  23. 23.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022. 8-bit optimizers via block-wise quantization. In ICLR.
  24. 24.T. Diggelmann, Jordan L. Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. Climate-fever: A dataset for verification of real-world climate claims. ArXiv.
  25. 25.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
  26. 26.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In NAACL.
  27. 27.Matthew Dunn, Levent Sagun, Mike Higgins, V. U. Guney, Volkan Cirik, and Kyunghyun Cho. 2017. Searchqa: A new q&a dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179.
  28. 28.Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018. T-REx: A large scale alignment of natural language with knowledge base triples. In LREC.
  29. 29.Theodoros Evgeniou and Massimiliano Pontil. 2004. Regularized multi–task learning. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining.
  30. 30.Manaal Faruqui and Dipanjan Das. 2018. Identifying well-formed natural language questions. In EMNLP.
  31. 31.Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML.
  32. 32.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007. The third pascal recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL workshop on textual entailment and paraphrasing.
  33. 33.Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2012. SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In The First Joint Conference on Lexical and Computational Semantics (SemEval).
  34. 34.Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo. 2012. Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports. Journal of Biomedical Informatics.
  35. 35.Luheng He, Mike Lewis, and Luke Zettlemoyer. 2015. Question-answer driven semantic role labeling: Using natural language to annotate natural language. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing.
  36. 36.Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Furstenau, Manfred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum. 2011. Robust disambiguation of named entities in text. In EMNLP.
  37. 37.Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn't always right. In EMNLP.
  38. 38.Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran. 2001. Toward semantics-based answer pinpointing. In Proceedings of the First International Conference on Human Language Technology Research.
  39. 39.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Cosmos QA: Machine reading comprehension with contextual commonsense reasoning. In EMNLP.
  40. 40.Kelvin Jiang, Dekun Wu, and Hui Jiang. 2019. FreebaseQA: A new factoid QA data set matching trivia-style question-answer pairs with Freebase. In NAACL-HLT.
  41. 41.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In NAACL-HLT.
  42. 42.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UnifiedQA: Crossing format boundaries with a single qa system. In Findings of EMNLP.
  43. 43.Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2019. QASC: A dataset for question answering via sentence composition. In AAAI.
  44. 44.Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2020. Qasc: A dataset for question answering via sentence composition. In AAAI.
  45. 45.Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. Scitail: A textual entailment dataset from science question answering. In AAAI.
  46. 46.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR.
  47. 47.Tomas Kocisky, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gabor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. TACL.
  48. 48.Neema Kotonya and Francesca Toni. 2020. Explainable automated fact-checking for public health claims. In EMNLP.
  49. 49.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. TACL.
  50. 50.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. 2017. RACE: Large-scale reading comprehension dataset from examinations. In EMNLP.
  51. 51.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, D. Kontokostas, Pablo N. Mendes, Sebastian Hellmann, M. Morsey, Patrick van Kleef, S. Auer, and C. Bizer. 2015. Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia. Semantic Web.
  52. 52.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning.
  53. 53.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero-shot relation extraction via reading comprehension. In CoNLL.
  54. 54.Xin Li and Dan Roth. 2002. Learning question classifiers. In COLING.
  55. 55.Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020. Birds have four legs?! NumerSense: Probing Numerical Commonsense Knowledge of Pre-Trained Language Models. In EMNLP.
  56. 56.Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019. Reasoning over paragraph effects in situations. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering.
  57. 57.Annie Louis, Dan Roth, and Filip Radlinski. 2020. ‘‘I'd rather just go to bed'': Understanding indirect answers. In EMNLP.
  58. 58.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  59. 59.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies.
  60. 60.Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala. 2014. Good debt or bad debt: Detecting semantic orientations in economic texts. J. Assoc. Inf. Sci. Technol.
  61. 61.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In LREC.
  62. 62.Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2020. Hatexplain: A benchmark dataset for explainable hate speech detection. arXiv preprint arXiv:2012.10289.
  63. 63.Julian McAuley and J. Leskovec. 2013. Hidden factors and hidden topics: understanding rating dimensions with review text. Proceedings of the 7th ACM conference on Recommender systems.
  64. 64.Clara H. McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain. 2020. Effective transfer learning for identifying similar questions: Matching user questions to covid-19 faqs. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.
  65. 65.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017. Mixed precision training. In ICLR.
  66. 66.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP.
  67. 67.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Noisy channel language model prompting for few-shot text classification. In ACL.
  68. 68.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021. Reframing instructional prompts to gptk's language. arXiv preprint arXiv:2109.07830.
  69. 69.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In ACL.
  70. 70.Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. 2020. Ethos: an online hate speech detection dataset. arXiv preprint arXiv:2006.08328.
  71. 71.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In EMNLP.
  72. 72.Courtney Napoles, Matthew Gormley, and Benjamin Van Durme. 2012. Annotated Gigaword. In Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction (AKBC-WEKEX).
  73. 73.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don't give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In EMNLP.
  74. 74.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In ACL.
  75. 75.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL.
  76. 76.Dimitris Pappas, Petros Stavropoulos, Ion Androutsopoulos, and Ryan McDonald. 2020. BioMRC: A dataset for biomedical machine reading comprehension. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing.
  77. 77.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS.
  78. 78.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In NeurIPS.
  79. 79.Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktaschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. How context affects language models' factual predictions. In Automated Knowledge Base Construction.
  80. 80.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In EMNLP.
  81. 81.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. In NAACL-HLT.
  82. 82.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  83. 83.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don't know: Unanswerable questions for squad. In ACL.
  84. 84.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In EMNLP.
  85. 85.Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013. Mctest: A challenge dataset for the open-domain machine comprehension of text. In EMNLP.
  86. 86.Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020. Getting closer to ai complete question answering: A set of prerequisite real tasks. In AAAI.
  87. 87.Sebastian Ruder. 2017. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098.
  88. 88.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020a. WINOGRANDE: an adversarial winograd schema challenge at scale. In AAAI.
  89. 89.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020b. Winogrande: An adversarial winograd schema challenge at scale. In AAAI.
  90. 90.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In ICLR.
  91. 91.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019a. Social IQa: Commonsense reasoning about social interactions. In EMNLP.
  92. 92.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019b. Social iqa: Commonsense reasoning about social interactions. In EMNLP-IJCNLP.
  93. 93.Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018. CARER: Contextualized affect representations for emotion recognition. In EMNLP.
  94. 94.Emily Sheng and David Uthus. 2020. Investigating societal biases in a poetry composition system. In Proceedings of the Second Workshop on Gender Bias in Natural Language Processing.
  95. 95.Damien Sileo, Tim Van De Cruys, Camille Pradel, and Philippe Muller. 2019. Mining discourse markers for unsupervised sentence representation learning. In NAACL-HLT.
  96. 96.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP.
  97. 97.Kai Sun, Dian Yu, Jianshu Chen, Dong Yu, Yejin Choi, and Claire Cardie. 2019. DREAM: A challenge data set and models for dialogue-based reading comprehension. TACL.
  98. 98.Oyvind Tafjord, Peter Clark, Matt Gardner, Wen-tau Yih, and Ashish Sabharwal. 2019a. Quarel: A dataset and models for answering questions about qualitative relationships. In AAAI.
  99. 99.Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019b. QuaRTz: An open-domain dataset of qualitative relationship questions. In EMNLP.
  100. 100.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In NAACL-HLT.
  101. 101.Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, and Antoine Bosselut. 2019. WIQA: A dataset for ‘‘what if...'' reasoning over procedural text. In EMNLP.
  102. 102.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In NAACL-HLT.
  103. 103.Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. Newsqa: A machine comprehension dataset. In Rep4NLP@ACL.
  104. 104.Sowmya Vajjala and Ivana Lucic. 2018. OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification. In Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications.
  105. 105.Ricardo Vilalta and Youssef Drissi. 2002. A perspective view and survey of meta-learning. Artificial intelligence review.
  106. 106.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. In BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP.
  107. 107.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  108. 108.William Yang Wang. 2017. ‘‘liar, liar pants on fire'': A new benchmark dataset for fake news detection. In ACL.
  109. 109.Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. Blimp: The benchmark of linguistic minimal pairs for english. TACL.
  110. 110.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. TACL.
  111. 111.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In ICLR.
  112. 112.Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy Usergenerated Text.
  113. 113.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL-HLT.
  114. 114.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In EMNLP: System Demonstrations.
  115. 115.Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni, Mo Yu, Shiyu Chang, Xiaoxiao Guo, and William Yang Wang. 2019. TWEETQA: A social media focused question answering dataset. In ACL.
  116. 116.Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A challenge dataset for open-domain question answering. In EMNLP.
  117. 117.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In EMNLP.
  118. 118.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. CrossFit: A few-shot learning challenge for cross-task generalization in nlp. In EMNLP.
  119. 119.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In EMNLP.
  120. 120.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. SWAG: A large-scale adversarial dataset for grounded commonsense inference. In EMNLP.
  121. 121.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In ACL.
  122. 122.Sheng Zhang, X. Liu, J. Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885.
  123. 123.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NeurIPS.
  124. 124.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In NAACL-HLT.
  125. 125.Tony Z Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In ICML.
  126. 126.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. 2021. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. In Findings of EMNLP.
  127. 127.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103.
  128. 128.Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth. 2019. ‘‘going on a vacation'' takes longer than ‘‘going for a walk'': A study of temporal commonsense understanding. In EMNLP.

Citation

MLA
Min, S., et al. “MetaICL: Learning to Learn In Context”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 2791–809, https://doi.org/10.18653/v1/2022.naacl-main.201.
APA
Min, S., Lewis, M., Zettlemoyer, L., & Hajishirzi, H. (2022). MetaICL: Learning to Learn In Context. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2791–2809. https://doi.org/10.18653/v1/2022.naacl-main.201
Chicago
Min, S., M. Lewis, L. Zettlemoyer, and H. Hajishirzi. 2022. “MetaICL: Learning to Learn In Context”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2791–2809. https://doi.org/10.18653/v1/2022.naacl-main.201.
Harvard
Min, S. et al. (2022) “MetaICL: Learning to Learn In Context”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 2791–2809. Available at: https://doi.org/10.18653/v1/2022.naacl-main.201.
Vancouver
1. Min S, Lewis M, Zettlemoyer L, Hajishirzi H (2022) MetaICL: Learning to Learn In Context. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 2791–2809

BibTeX

@inproceedings{min-etal-2022-metaicl,
    title = "{M}eta{ICL}: Learning to Learn In Context",
    author = "Min, Sewon  and
      Lewis, Mike  and
      Zettlemoyer, Luke  and
      Hajishirzi, Hannaneh",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.201/",
    doi = "10.18653/v1/2022.naacl-main.201",
    pages = "2791--2809"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/