Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data
Shuohang WangYichong XuYuwei FangYang LiuSiqi SunRuochen XuChenguang ZhuMichael Zeng
Proposes REINA, a lightweight supervised method that significantly improves natural language understanding and generation performance by retrieving and concatenating similar labeled training examples with the input text rather than searching costly external corpora.
Modern natural language processing systems increasingly rely on expanding model sizes or searching massive external text collections to improve accuracy. However, maintaining and querying massive external databases incurs substantial computational overhead and drastically slows down inference speeds. Meanwhile, even large models with hundreds of millions of parameters struggle to retain every pattern present in their original supervised training data.
The article introduces and evaluates REINA (Retrieving from the Training Data), an approach that improves model performance by retrieving relevant labeled examples directly from a task's own training data during both training and evaluation. The primary objective is to demonstrate that retrieving from existing supervised datasets offers a computationally efficient alternative to querying massive external corpora or scaling up model parameters.
To assess the method, the researchers conducted extensive empirical evaluations across four standard language task domains—text summarization, language modeling, machine translation, and question answering—spanning 12 benchmark datasets. The framework uses a standard, fast BM25 ranking algorithm to index training pairs. When processing an input, the system retrieves the most relevant training instances, filters out exact duplicates during training to prevent data leakage, concatenates the retrieved content with the input query, and feeds the combined sequence into standard transformer models.
The evaluation yielded several key findings. First, integrating REINA improved baseline performance across 11 of the 12 evaluated datasets, achieving state-of-the-art results on benchmark datasets including XSum, BigPatent, and CommonsenseQA (reaching first place on the public leaderboard). Second, the approach enabled smaller models to outperform larger architectures; for example, a base summarization model using REINA outperformed models with more than twice the parameter count on BigPatent and WikiHow. Third, the system scaled effectively when drawing from external training collections via retrieval, outperforming models directly trained on merged data pools. Finally, while performance increased on larger corpora, gains did not materialize on very small training sets like WikiText2, where retrieval pool diversity was limited.
These results demonstrate that providing closely matching training examples as in-context reminders allows language models to recall critical task-specific information more reliably. For organizations deploying machine learning systems, this approach reduces computational costs and infrastructure requirements, as smaller models paired with simple training data retrieval can match or exceed the accuracy of much larger, more expensive models.
Organizations developing language processing pipelines should consider indexing their existing supervised datasets as an immediate, low-cost enhancement before scaling model sizes. Where applicable, integrating external structured knowledge can further boost retrieval quality for complex reasoning tasks. However, practitioners should note that the method relies on a sufficiently large and diverse training corpus to locate relevant examples. Overall confidence in the reported improvements is high across standard supervised tasks, though cautious evaluation is recommended when working with highly constrained or small-scale datasets.
- Paper: Generalization through Memorization: Nearest Neighbor Language Models, Urvashi Khandelwal et al. (2020). Introduces nearest-neighbor language modeling over stored training datastores, establishing the foundational principle of retrieving from memorized data during prediction that REINA adapts to supervised contexts.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). Demonstrates the effectiveness of retrieving semantically similar training examples to form in-context prompts for language models, providing core motivation for REINA's in-context training retrieval paradigm.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Establishes the foundational retrieval-augmented generation framework, which REINA simplifies by querying task-specific training data rather than massive external knowledge corpora.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Provides the conceptual groundwork for integrating neural text retrieval into language modeling objectives during both pre-training and inference.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). Presents Fusion-in-Decoder architecture for conditioning sequence generation on retrieved passages, serving as a primary baseline and structural reference for retrieval-augmented sequence models.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Introduces the standard sequence-to-sequence BART architecture used as the primary generative backbone across multiple benchmark tasks evaluated in REINA.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). Extends the concept of simple training-free retrieval concatenation directly to frozen, off-the-shelf autoregressive language models across diverse scales.
- Paper: REPLUG: Retrieval-Augmented Black-Box Language Models, Weijia Shi et al. (2024). Generalizes in-context retrieval augmentation to black-box commercial language models accessed exclusively via inference APIs without internal parameter fine-tuning.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Advances basic retrieval augmentation into an adaptive, self-reflective framework where models dynamically determine when retrieval is needed and critique retrieved context quality.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). Builds upon single-step retrieval approaches by introducing dynamic, forward-looking active retrieval during multi-sentence generation.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Directly tackles the vulnerability of retrieval-augmented language models to irrelevant or misleading retrieved context through adaptive adversarial training.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). Provides a comprehensive taxonomy and survey of how modern retrieval-augmented generation paradigms evolved and scaled across large language models.
