Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data

Shuohang WangYichong XuYuwei FangYang LiuSiqi SunRuochen XuChenguang ZhuMichael Zeng

article2022ACL117 citations

Proposes REINA, a lightweight supervised method that significantly improves natural language understanding and generation performance by retrieving and concatenating similar labeled training examples with the input text rather than searching costly external corpora.

Listen

Modern natural language processing systems increasingly rely on expanding model sizes or searching massive external text collections to improve accuracy. However, maintaining and querying massive external databases incurs substantial computational overhead and drastically slows down inference speeds. Meanwhile, even large models with hundreds of millions of parameters struggle to retain every pattern present in their original supervised training data.

The article introduces and evaluates REINA (Retrieving from the Training Data), an approach that improves model performance by retrieving relevant labeled examples directly from a task's own training data during both training and evaluation. The primary objective is to demonstrate that retrieving from existing supervised datasets offers a computationally efficient alternative to querying massive external corpora or scaling up model parameters.

To assess the method, the researchers conducted extensive empirical evaluations across four standard language task domains—text summarization, language modeling, machine translation, and question answering—spanning 12 benchmark datasets. The framework uses a standard, fast BM25 ranking algorithm to index training pairs. When processing an input, the system retrieves the most relevant training instances, filters out exact duplicates during training to prevent data leakage, concatenates the retrieved content with the input query, and feeds the combined sequence into standard transformer models.

The evaluation yielded several key findings. First, integrating REINA improved baseline performance across 11 of the 12 evaluated datasets, achieving state-of-the-art results on benchmark datasets including XSum, BigPatent, and CommonsenseQA (reaching first place on the public leaderboard). Second, the approach enabled smaller models to outperform larger architectures; for example, a base summarization model using REINA outperformed models with more than twice the parameter count on BigPatent and WikiHow. Third, the system scaled effectively when drawing from external training collections via retrieval, outperforming models directly trained on merged data pools. Finally, while performance increased on larger corpora, gains did not materialize on very small training sets like WikiText2, where retrieval pool diversity was limited.

These results demonstrate that providing closely matching training examples as in-context reminders allows language models to recall critical task-specific information more reliably. For organizations deploying machine learning systems, this approach reduces computational costs and infrastructure requirements, as smaller models paired with simple training data retrieval can match or exceed the accuracy of much larger, more expensive models.

Organizations developing language processing pipelines should consider indexing their existing supervised datasets as an immediate, low-cost enhancement before scaling model sizes. Where applicable, integrating external structured knowledge can further boost retrieval quality for complex reasoning tasks. However, practitioners should note that the method relies on a sufficiently large and diverse training corpus to locate relevant examples. Overall confidence in the reported improvements is high across standard supervised tasks, though cautious evaluation is recommended when working with highly constrained or small-scale datasets.

arXiv: 2203.08773microsoft/REINA
Cover for Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data

Abstract

Retrieval-based methods have been shown to be effective in NLP tasks via introducing external knowledge. However, the indexing and retrieving of large-scale corpora bring considerable computational cost. Surprisingly, we found that REtrieving from the traINing datA (REINA) only can lead to significant gains on multiple NLG and NLU tasks. We retrieve the labeled training instances most similar to the input text and then concatenate them with the input to feed into the model to generate the output. Experimental results show that this simple method can achieve significantly better performance on a variety of NLU and NLG tasks, including summarization, machine translation, language modeling, and question answering tasks. For instance, our proposed method achieved state-of-the-art results on XSum, BigPatent, and CommonsenseQA. Our code is released.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Retrieval-based Methods
  • 3.2 Retrieval from Training Data ( REINA )
  • 3.3 Model Training and Inference
  • 4 Experiment
  • 4.1 Dataset
  • 4.2 REINA Details
  • 4.3 Optimization Details
  • 4.4 Experiment Results
  • 4.5 Further Analysis
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — REINA Supervised Retrieval-Augmented Framework

    model/method

    REtrieving from the traINing datA (REINA) is a supervised learning framework that enhances sequence modeling by retrieving labeled training instances similar to the input and concatenating them directly into the input sequence during both training and inference.

    Given a supervised training dataset T={(x1,y1),(x2,y2),…,(xN,yN)}T = \{(x_1, y_1), (x_2, y_2), \dots, (x_N, y_N)\}, where xix_i is an input text and yiy_i is its ground-truth label, REINA uses TT directly as an indexed corpus C={(ki,vi)}i=1N\mathcal{C} = \{(k_i, v_i)\}_{i=1}^N of key-value pairs.

    For any query sequence xx, a retrieval engine E\mathcal{E} based on BM25 sparse matching returns the top-KK most relevant entries:

    {(ki1,vi1),…,(kiK,viK)}=E(x∣C)\{(k_{i_1}, v_{i_1}), \dots, (k_{i_K}, v_{i_K})\} = \mathcal{E}(x \mid \mathcal{C})

    A task-specific combination function ff combines the input xx with the retrieved keys and/or values, and feeds the resulting sequence into a neural model M\mathcal{M} (such as a sequence-to-sequence or encoder-only Transformer) to produce the prediction OO:

    O=M(f(x,{(ki1,vi1),…,(kiK,viK)}))O = \mathcal{M}\left(f(x, \{(k_{i_1}, v_{i_1}), \dots, (k_{i_K}, v_{i_K})\})\right)

    During model training, the query instance itself is filtered out from the retrieved results to prevent label leakage. During test inference, no filtering is performed, and candidate pairs are retrieved across the full training dataset.

  2. Knowl 2 — Task-Specific Formulations for Generation Tasks in REINA

    model/method

    REINA customizes key-value indexing and combination functions ff across different natural language generation tasks:

    1. Abstractive Summarization:

      • Key kik_i: Source document xix_i.
      • Value viv_i: Target summary yiy_i.
      • Retrieval Query: Input document xx.
      • Combination Function: Because source documents are long, only the retrieved target summaries are retained and appended to the input: fsumm(x,{(xij,yij)}j=1K)=[x;yi1;… ;yiK]f_{\text{summ}}(x, \{(x_{i_j}, y_{i_j})\}_{j=1}^K) = [x; y_{i_1}; \dots; y_{i_K}]
    2. Seq2Seq Chunk-Level Language Modeling:

      • Corpus text is segmented into fixed-size contiguous chunks {C1,C2,… }\{C_1, C_2, \dots\}.
      • Key kik_i: Context chunk CiC_i.
      • Value viv_i: Subsequent chunk Ci+1C_{i+1}.
      • Retrieval Query: Context chunk xx.
      • Combination Function: Retrieved subsequent chunks are prepended to the context chunk xx: fLM(x,{(xij,yij)}j=1K)=[yi1;… ;yiK;x]f_{\text{LM}}(x, \{(x_{i_j}, y_{i_j})\}_{j=1}^K) = [y_{i_1}; \dots; y_{i_K}; x] The Seq2Seq model then autoregressively generates the target next chunk.
    3. Machine Translation:

      • Key kik_i: Source-language sentence xi∈Sx_i \in S.
      • Value viv_i: Target-language translation yi∈Ty_i \in T.
      • Retrieval Query: Source sentence xx.
      • Combination Function: Retrieved target translations are appended to the input source sentence: fMT(x,{(xij,yij)}j=1K)=[x;yi1;… ;yiK]f_{\text{MT}}(x, \{(x_{i_j}, y_{i_j})\}_{j=1}^K) = [x; y_{i_1}; \dots; y_{i_K}]
  3. Knowl 3 — Knowledge-Augmented Multiple-Choice Question Answering in REINA

    model/method

    For multiple-choice question answering, each instance comprises a question xx, a set of candidate answer choices {c1,…,cM}\{c_1, \dots, c_M\}, and a ground-truth choice yy.

    In standard REINA for QA, the index contains pairs where each key is the concatenated string of question xix_i and correct answer yiy_i, and the value is the pair (xi,yi)(x_i, y_i). For a candidate choice cmc_m, the query is [x;cm][x; c_m], and the combination function concatenates both questions and answers with the query:

    fQA((x,cm),{(xij,yij)}j=1K)=[x;cm;xi1;yi1;… ;xiK;yiK]f_{\text{QA}}((x, c_m), \{(x_{i_j}, y_{i_j})\}_{j=1}^K) = [x; c_m; x_{i_1}; y_{i_1}; \dots; x_{i_K}; y_{i_K}]

    To overcome limited lexical overlap in commonsense reasoning, REINA introduces a knowledge-augmented index. For each question xx and choice cc, string matching identifies ConceptNet entity sets E(x)E(x) and E(c)E(c). The longest matched concepts are selected, and their definitions are fetched from Wiktionary. Connecting relational triples R={(e1,r,e2)∣e1∈E(x),e2∈E(c),(e1,e2)∈ConceptNet}R = \{(e_1, r, e_2) \mid e_1 \in E(x), e_2 \in E(c), (e_1, e_2) \in \text{ConceptNet}\} are collected. Definitions and relational triples are concatenated to create a knowledge context KK.

    The query [x;cm;K][x; c_m; K] is then matched against indexed training entries [xi;yi;Ki][x_i; y_i; K_i], and retrieved exemplars are concatenated for scoring by an encoder-only Transformer.

  4. Knowl 4 — Performance of REINA on Abstractive Summarization Benchmarks

    data/table

    REINA was evaluated across five summarization datasets spanning single-document news (XSum), multi-document news (Multi-News), broad multi-domain news (NEWSROOM), how-to procedural instructions (WikiHow), and technical patent documents (BigPatent). Models were initialized with PEGASUS-large, BART-base, and BART-large, and evaluated via ROUGE-1 (R-1), ROUGE-2 (R-2), and ROUGE-L (R-L).

    BigPatent XSum WikiHow Multi-News NEWSROOM
    Model R-1 R-2 R-L R-1 R-2 R-L R-1 R-2 R-L R-1 R-2 R-L R-1 R-2 R-L
    Earlier SOTA 37.5 10.6 22.7 45.1 22.2 37.2 28.5 9.2 26.5 43.4 14.8 17.4 39.9 28.3 36.8
    PEGASUS (paper) 53.6 33.2 42.3 47.2 24.6 39.3 43.1 19.7 34.8 47.5 18.7 24.9 45.2 33.5 41.3
    PEGASUS 38.4 13.5 26.3 46.6 23.9 38.6 35.9 15.3 30.3 43.1 15.4 22.6 41.7 30.7 37.8
    REINA (PG) 44.6 21.5 33.0 48.2 26.0 40.2 36.8 16.7 31.0 45.0 17.1 23.8 41.4 30.5 37.5
    BART-base 44.2 16.9 28.4 41.0 18.2 33.3 43.3 18.1 33.9 44.8 16.4 23.3 41.3 29.1 37.5
    REINA (B) 59.5 42.6 50.6 43.2 21.0 35.5 44.2 19.4 34.9 45.1 16.9 23.6 41.2 29.0 37.5
    BART-large 44.9 17.5 28.9 44.7 21.6 36.5 43.4 19.0 34.9 44.1 16.6 22.7 41.6 29.4 38.0
    REINA (L) 60.7 43.3 51.3 46.5 24.1 38.6 44.2 20.4 35.8 46.9 17.7 24.0 42.5 30.2 38.7

    REINA consistently improves over corresponding baseline initializations across all five datasets. REINA with PEGASUS achieves state-of-the-art performance on XSum (48.2 R-1, 26.0 R-2, 40.2 R-L), while REINA with BART-large sets the state-of-the-art on BigPatent (60.7 R-1, 43.3 R-2, 51.3 R-L). Notably, REINA with BART-base (59.5 R-1 on BigPatent, 44.2 R-1 on WikiHow) outperforms vanilla BART-large (44.9 R-1 on BigPatent, 43.4 R-1 on WikiHow), demonstrating that retrieval from training data enables smaller models to rival or outperform models with twice the parameters.

  5. Knowl 5 — Scaling Training Data via Cross-Corpus Retrieval in Summarization

    data/table

    The impact of expanding the retrieval index using auxiliary external corpora was evaluated on the XSum test set using BART-large. In multi-task joint training, the model was trained on the combined data. In REINA, the model was trained strictly on XSum document-summary pairs, while the BM25 retrieval index was constructed from external datasets: CNN/DailyMail (CNN), NEWSROOM (NR), and the 160GB BART pre-training corpus (160G).

    Model R-1 R-2 R-L
    BART (XSum) 44.7 21.6 36.5
    BART (XSum+CNN) 44.6 21.6 36.9
    REINA (XSum) 46.5 24.1 38.6
    REINA (XSum+CNN) 47.5 25.2 39.5
    REINA (XSum+NR) 47.5 24.9 39.4
    REINA (XSum+160G) 47.7 25.1 39.5

    Jointly training on merged datasets (XSum + CNN) yields no improvement over standard training on XSum alone (44.6 vs. 44.7 R-1). Conversely, expanding REINA's retrieval index with CNN/DailyMail, NEWSROOM, or the 160GB pretraining corpus increases performance by approximately 1.0 R-1 point (up to 47.7 R-1) without modifying the supervised fine-tuning target.

  6. Knowl 6 — Performance of REINA on Commonsense Question Answering Benchmarks

    data/table

    REINA was evaluated on CommonsenseQA (CSQA), Abductive NLI (aNLI), and Physical IQA (PIQA) using DeBERTa (1.5B parameters), comparing vanilla retrieval against knowledge-augmented retrieval (incorporating ConceptNet and Wiktionary).

    Model CSQA aNLI PIQA
    Dev Set
    DeBERTa 84.0 88.8 85.6
    REINA (w/o K) 88.8 88.6 85.5
    REINA (w/ K) 86.8 89.6 86.9
    Test Set
    CALM 71.8 82.4 76.9
    UNICORN 79.3 87.3 90.1
    DEKCOR 83.3 - -
    DeBERTa - 86.8 85.1
    REINA 84.6 88.0 85.4

    REINA achieved state-of-the-art performance on the CSQA hidden test leaderboard (84.6% accuracy). On the dev sets, direct retrieval without external knowledge (w/o K) performed best for CSQA (88.8%), whereas integrating ConceptNet and Wiktionary knowledge (w/ K) was beneficial for aNLI (89.6% vs. 88.6%) and PIQA (86.9% vs. 85.5%).

  7. Knowl 7 — Language Modeling with Chunk-Level Training Set Retrieval

    data/table

    REINA was evaluated on language modeling using WikiText103 and WikiText2 in a chunk-level Seq2Seq framework (chunk size of 128 words for WikiText103 and 64 words for WikiText2), evaluated by Perplexity (PPL, lower is better).

    Model WikiText103 WikiText2
    Transformer-XL 18.30 -
    kNN-LM 15.79 -
    GPT-2 17.48 18.34
    BART-Base 15.88 20.41
    REINA (B) 14.76 20.78
    BART-Large 12.10 15.11
    REINA (L) 11.36 15.62

    On WikiText103, REINA with BART-Large achieves a perplexity of 11.36, outperforming standard BART-Large (12.10), kNN-LM (15.79), and GPT-2 (17.48). On WikiText2, REINA does not yield gains over baselines (15.62 vs. 15.11 PPL for BART-Large), indicating that REINA's sparse retrieval requires a sufficiently large index to surface useful related training contexts.

  8. Knowl 8 — Performance of REINA on Neural Machine Translation

    data/table

    REINA was evaluated on machine translation using WMT16 English-Turkish (low-resource) and English-German (high-resource) in both translation directions, initialized with mBART-large and evaluated using SacreBLEU.

    Model en→\totr tr→\toen en→\tode de→\toen
    XLM - - 26.4 34.3
    mBART 18.4 23.1 32.6 37.0
    REINA 18.8 23.6 32.9 37.0

    REINA improves translation quality over the strong mBART baseline across en→\totr (18.8 vs. 18.4), tr→\toen (23.6 vs. 23.1), and en→\tode (32.9 vs. 32.6), while matching mBART on de→\toen (37.0).

  9. Knowl 9 — Architectural and Computational Characteristics of REINA

    model/method

    REINA differs from existing retrieval-augmented and nearest-neighbor generation methods in three primary dimensions:

    1. Sparse Indexing over Dense Passage Retrieval (DPR): REINA indexes the training set with BM25 via Lucene. This bypasses the need for training task-specific neural encoders with parallel data and eliminates DPR's document length restriction (typically ≤128\le 128 tokens).
    2. Chunk/Document-Level Retrieval vs. Token-Level kNN: In contrast to kNN-LM and kNN-MT, which perform dense nearest-neighbor searches in hidden state space at every decoding step (incurring a generation slowdown of up to two orders of magnitude), REINA executes a single sparse retrieval per chunk or document prior to sequence encoding.
    3. Supervised Parameter Fine-Tuning vs. In-Context Few-Shot Prompting: Unlike few-shot in-context learning which relies solely on prompt exemplars with frozen models, REINA fine-tunes all model parameters on concatenated input-exemplar sequences across the full training dataset.

Coverage note — Qualitative case study examples from Table 7 were omitted as they serve solely to illustrate individual retrieval outputs without introducing distinct empirical benchmarks or methods.

References

  1. 1.Chenxin An, Ming Zhong, Zhichao Geng, Jianqiang Yang, and Xipeng Qiu. 2021. Retrievalsum: A retrieval enhanced framework for abstractive summarization. arXiv preprint arXiv:2109.07943.
  2. 2.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. International Conference on Learning Representations (ICLR).
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In AAAI Conference on Artificial Intelligence (AAAI).
  4. 4.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, et al. 2016. Findings of the 2016 conference on machine translation. In First Conference on Machine Translation.
  5. 5.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Neural Information Processing Systems (NeurIPS).
  6. 6.Deng Cai, Yan Wang, Huayang Li, Wai Lam, and Lemao Liu. 2021. Neural machine translation with monolingual translation memory. Association for Computational Linguistics (ACL).
  7. 7.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer opendomain questions. Association for Computational Linguistics (ACL).
  8. 8.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a fixed-length context. Association for Computational Linguistics (ACL).
  9. 9.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2020. Unified language model pre-training for natural language understanding and generation. International Conference on Machine Learning (ICML).
  10. 10.Alexander R Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R Radev. 2019. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. Association for Computational Linguistics (ACL).
  11. 11.Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. North American Chapter of the Association for Computational Linguistics (NAACL).
  12. 12.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2018. Search engine guided neural machine translation. In AAAI Conference on Artificial Intelligence (AAAI).
  13. 13.Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang. 2018. Generating sentences by editing prototypes. Transactions of the Association for Computational Linguistics (TACL).
  14. 14.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrievalaugmented language model pre-training. International Conference on Machine Learning (ICML).
  15. 15.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: Decoding-enhanced bert with disentangled attention. International Conference on Learning Representations (ICLR).
  16. 16.Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for opendomain question answering. Empirical Methods in Natural Language Processing (EMNLP).
  17. 17.Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2021. Nearest neighbor machine translation. International Conference on Learning Representations (ICLR).
  18. 18.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. International Conference on Learning Representations (ICLR).
  19. 19.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. International Conference on Learning Representations (ICLR).
  20. 20.Mahnaz Koupaee and William Yang Wang. 2018. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305.
  21. 21.Guillaume Lample and Alexis Conneau. 2019. Crosslingual language model pretraining. Neural Information Processing Systems (NeurIPS).
  22. 22.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. Association for Computational Linguistics (ACL).
  23. 23.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. Paq: 65 million probably-asked questions and what you can do with them. Transactions of the Association for Computational Linguistics.
  24. 24.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. Association for Computational Linguistics (ACL).
  25. 25.Yankai Lin, Haozhe Ji, Zhiyuan Liu, and Maosong Sun. 2018. Denoising distantly supervised open-domain question answering. In Association for Computational Linguistics (ACL).
  26. 26.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804.
  27. 27.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pretraining for neural machine translation. Transactions of the Association for Computational Linguistics (TACL).
  28. 28.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. International Conference on Learning Representations (ICLR).
  29. 29.Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark. AAAI Conference on Artificial Intelligence (AAAI).
  30. 30.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. International Conference on Learning Representations (ICLR).
  31. 31.Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016. Abstractive text summarization using sequence-to-sequence rnns and beyond. SIGNLL Conference on Computational Natural Language Learning (CoNLL).
  32. 32.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. Empirical Methods in Natural Language Processing (EMNLP).
  33. 33.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2030. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR).
  35. 35.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Now Publishers Inc.
  36. 36.Hinrich Schütze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval. Cambridge University Press Cambridge.
  37. 37.Eva Sharma, Chen Li, and Lu Wang. 2019. Bigpatent: A large-scale dataset for abstractive and coherent summarization. Association for Computational Linguistics (ACL).
  38. 38.Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. Conceptnet 5.5: An open multilingual graph of general knowledge. In AAAI conference on artificial intelligence (AAAI).
  39. 39.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. North American Chapter of the Association for Computational Linguistics (NAACL).
  40. 40.Shuohang Wang, Mo Yu, Xiaoxiao Guo, Zhiguo Wang, Tim Klinger, Wei Zhang, Shiyu Chang, Gerry Tesauro, Bowen Zhou, and Jing Jiang. 2018. R 3: Reinforced ranker-reader for open-domain question answering. In AAAI Conference on Artificial Intelligence (AAAI).
  41. 41.Jiacheng Xu and Greg Durrett. 2021. Dissecting generation modes for abstractive summarization models via ablation and attribution. Association for Computational Linguistics (ACL).
  42. 42.Yichong Xu, Chenguang Zhu, Ruochen Xu, Yang Liu, Michael Zeng, and Xuedong Huang. 2021. Fusing context into knowledge graph for commonsense question answering. In Association for Computational Linguistics (ACL).
  43. 43.Donghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, and Michael Zeng. 2021. Kg-fid: Infusing knowledge graph in fusion-in-decoder for open-domain question answering. arXiv preprint arXiv:2110.04330.
  44. 44.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gapsentences for abstractive summarization. In International Conference on Machine Learning (ICML).
  45. 45.Yizhe Zhang, Siqi Sun, Xiang Gao, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2021. Joint retrieval and generation training for grounded text generation. arXiv preprint arXiv:2105.06597.
  46. 46.Wangchunshu Zhou, Dong-Ho Lee, Ravi Kiran Selvam, Seyeon Lee, and Xiang Ren. 2021. Pre-training text-to-text transformers for concept-centric common sense. In International Conference on Learning Representations (ICLR).

Citation

MLA
Wang, S., et al. “Training Data Is More Valuable Than You Think: A Simple and Effective Method by Retrieving from Training Data”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3170–79, https://doi.org/10.18653/v1/2022.acl-long.226.
APA
Wang, S., Xu, Y., Fang, Y., Liu, Y., Sun, S., Xu, R., Zhu, C., & Zeng, M. (2022). Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3170–3179. https://doi.org/10.18653/v1/2022.acl-long.226
Chicago
Wang, S., Y. Xu, Y. Fang, et al. 2022. “Training Data Is More Valuable Than You Think: A Simple and Effective Method by Retrieving from Training Data”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3170–79. https://doi.org/10.18653/v1/2022.acl-long.226.
Harvard
Wang, S. et al. (2022) “Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3170–3179. Available at: https://doi.org/10.18653/v1/2022.acl-long.226.
Vancouver
1. Wang S, Xu Y, Fang Y, Liu Y, Sun S, Xu R, Zhu C, Zeng M (2022) Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3170–3179

BibTeX

@inproceedings{wang-etal-2022-training,
    title = "Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data",
    author = "Wang, Shuohang  and
      Xu, Yichong  and
      Fang, Yuwei  and
      Liu, Yang  and
      Sun, Siqi  and
      Xu, Ruochen  and
      Zhu, Chenguang  and
      Zeng, Michael",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.226/",
    doi = "10.18653/v1/2022.acl-long.226",
    pages = "3170--3179"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/