Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning
Xiang ChenLei LiNingyu ZhangXiaozhuan LiangShumin DengChuanqi TanFei HuangLuo SiHuajun Chen
Proposes RETROPROMPT, a retrieval-augmented prompt learning framework that decouples knowledge from rote memorization using an open-book key-value datastore to improve generalization in few-shot and zero-shot NLP tasks.
Modern natural language processing models frequently struggle when adapting to new domains or learning from limited training examples. When trained with conventional prompt learning methods, these models often rely on rote memorization of atypical instances or overfit to shallow patterns, leading to unstable performance on new tasks.
The article introduces and evaluates RETROPROMPT, a retrieval-augmented framework designed to decouple factual knowledge from rote parameter memorization. The core objective is to improve model accuracy and stability across few-shot, zero-shot, fully-supervised, and cross-domain settings by providing an external reference system built directly from training data.
To accomplish this, the authors developed an open-book knowledge store that embeds training examples as key-value pairs using the model's contextual representations. RETROPROMPT integrates this retrieved information at three distinct stages: inserting aggregated neural demonstrations directly into the input embedding layer, using nearest-neighbor predictions to focus the loss function on hard instances during training, and interpolating non-parametric retrieval predictions with standard model outputs at inference time. The framework was evaluated across nine benchmark datasets spanning single-sentence classification, sentence-pair classification, and complex multi-class information extraction.
The evaluation revealed several key findings. First, RETROPROMPT consistently outperformed leading prompt-tuning baselines, achieving an average 16-shot accuracy of 75.6% across nine tasks compared to 71.4% for baseline prompt models. Second, the system demonstrated superior transferability to new domains, such as improving cross-domain accuracy from 20.9% to 49.4% when evaluated on transfer between paraphrase datasets. Third, the framework proved effective in zero-shot tasks (52.5% average accuracy versus 41.1% to 47.0% for baselines) and fully-supervised long-tail distributions by referencing unlabeled or stored examples without requiring external knowledge bases. Finally, influence-function analysis demonstrated that the system reduced the model's reliance on rote memorization, yielding lower memorization scores (0.032 versus 0.121 for standard prompts and 4.597 for fine-tuning) while decreasing performance variance across random seeds.
These findings indicate that integrating open-book retrieval directly into the training and inference pipeline provides a cost-effective, scalable way to improve model robustness without expanding overall parameter counts. Using compact neural demonstrations circumvents input sequence length bottlenecks, allowing the method to scale effectively to multi-class classification and information extraction tasks where standard discrete demonstrations fail.
Organizations developing low-resource natural language applications should consider adopting retrieval-augmented prompting architectures to enhance accuracy and reduce prediction instability. Prior to production deployment, engineering teams should conduct pilot evaluations to assess retrieval query latency and explore applying the architecture to generative and question-answering workloads.
While the empirical results are robust across the evaluated classification benchmarks, the current evidence is limited to natural language understanding tasks and relies on periodic asynchronous index refreshing. Further testing is necessary to confirm retrieval efficiency and performance on massive web-scale corpora and open-ended generative applications.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). Introduces prompt-based fine-tuning and nearest-neighbor demonstration sampling for few-shot learning, establishing the core paradigm that RetroPrompt extends.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). Demonstrates the benefit of retrieving semantically similar in-context examples for prompt conditioning, a direct prerequisite to RetroPrompt's neural demonstration framework.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Establishes foundational retrieval-augmented generation architectures that decouple parametric memory from external knowledge access.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Pioneers continuous soft prompt tuning with frozen backbones, which provides the underlying parameter-efficient framework used in RetroPrompt.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Introduces continuous prompt optimization (P-Tuning) to overcome the brittleness and memorization tendencies of discrete prompts.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive taxonomy and formal foundations for prompting methods and context demonstration techniques in NLP.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). Formulates cloze-style prompt engineering and pattern-exploiting training that underpins modern few-shot prompt adaptation.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Investigates how prompts elicit factual knowledge stored in pre-trained models versus relying on rote memorization.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Analyzes prediction instability and surface biases in few-shot prompting that motivate non-parametric retrieval calibration.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). Develops dynamic prompt retrieval from an external pool to adapt frozen models without catastrophic forgetting.
- Paper: UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation, Daixuan Cheng et al. (2023). Extends retrieval-augmented prompting by training a universal cross-task retriever to automatically fetch in-context prompts for zero-shot inputs.
- Paper: Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data, Shuohang Wang et al. (2022). Generalizes the concept of utilizing training-data retrieval directly at inference time across a broader suite of generative NLP tasks.
- Paper: REPLUG: Retrieval-Augmented Black-Box Language Models, Weijia Shi et al. (2024). Applies external retrieval integration to black-box language models while tuning the retriever using language model feedback.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Investigates the boundary between parametric memory and external non-parametric retrieval on long-tail factual knowledge.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). Builds on retrieval-augmented LLM architectures by unifying passage ranking and generation into a single instruction-tuned framework.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). Surveys the advanced landscape of modular retrieval-augmented generation pipelines developed in the wake of retrieval-augmented prompting methods.
- Paper: Memory-assisted prompt editing to improve GPT-3 after deployment, Aman Madaan et al. (2022). Applies dynamic key-value memory retrieval to prompt editing for post-deployment error correction.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). Adapts retrieval-augmented modeling into an off-the-shelf, in-context paradigm without modifying internal neural weights.
