Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning

Xiang ChenLei LiNingyu ZhangXiaozhuan LiangShumin DengChuanqi TanFei HuangLuo SiHuajun Chen

article2022NeurIPS65 citations

Proposes RETROPROMPT, a retrieval-augmented prompt learning framework that decouples knowledge from rote memorization using an open-book key-value datastore to improve generalization in few-shot and zero-shot NLP tasks.

Listen

Modern natural language processing models frequently struggle when adapting to new domains or learning from limited training examples. When trained with conventional prompt learning methods, these models often rely on rote memorization of atypical instances or overfit to shallow patterns, leading to unstable performance on new tasks.

The article introduces and evaluates RETROPROMPT, a retrieval-augmented framework designed to decouple factual knowledge from rote parameter memorization. The core objective is to improve model accuracy and stability across few-shot, zero-shot, fully-supervised, and cross-domain settings by providing an external reference system built directly from training data.

To accomplish this, the authors developed an open-book knowledge store that embeds training examples as key-value pairs using the model's contextual representations. RETROPROMPT integrates this retrieved information at three distinct stages: inserting aggregated neural demonstrations directly into the input embedding layer, using nearest-neighbor predictions to focus the loss function on hard instances during training, and interpolating non-parametric retrieval predictions with standard model outputs at inference time. The framework was evaluated across nine benchmark datasets spanning single-sentence classification, sentence-pair classification, and complex multi-class information extraction.

The evaluation revealed several key findings. First, RETROPROMPT consistently outperformed leading prompt-tuning baselines, achieving an average 16-shot accuracy of 75.6% across nine tasks compared to 71.4% for baseline prompt models. Second, the system demonstrated superior transferability to new domains, such as improving cross-domain accuracy from 20.9% to 49.4% when evaluated on transfer between paraphrase datasets. Third, the framework proved effective in zero-shot tasks (52.5% average accuracy versus 41.1% to 47.0% for baselines) and fully-supervised long-tail distributions by referencing unlabeled or stored examples without requiring external knowledge bases. Finally, influence-function analysis demonstrated that the system reduced the model's reliance on rote memorization, yielding lower memorization scores (0.032 versus 0.121 for standard prompts and 4.597 for fine-tuning) while decreasing performance variance across random seeds.

These findings indicate that integrating open-book retrieval directly into the training and inference pipeline provides a cost-effective, scalable way to improve model robustness without expanding overall parameter counts. Using compact neural demonstrations circumvents input sequence length bottlenecks, allowing the method to scale effectively to multi-class classification and information extraction tasks where standard discrete demonstrations fail.

Organizations developing low-resource natural language applications should consider adopting retrieval-augmented prompting architectures to enhance accuracy and reduce prediction instability. Prior to production deployment, engineering teams should conduct pilot evaluations to assess retrieval query latency and explore applying the architecture to generative and question-answering workloads.

While the empirical results are robust across the evaluated classification benchmarks, the current evidence is limited to natural language understanding tasks and relies on periodic asynchronous index refreshing. Further testing is necessary to confirm retrieval efficiency and performance on massive web-scale corpora and open-ended generative applications.

Cover for Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning

Abstract

Prompt learning approaches have made waves in natural language processing by inducing better few-shot performance while they still follow a parametric-based learning paradigm; the oblivion and rote memorization problems in learning may encounter unstable generalization issues. Specifically, vanilla prompt learning may struggle to utilize atypical instances by rote during fully-supervised training or over-fit shallow patterns with low-shot data. To alleviate such limitations, we develop RETROPROMPT with the motivation of decoupling knowledge from memorization to help the model strike a balance between generalization and memorization. In contrast with vanilla prompt learning, RETROPROMPT constructs an open-book knowledge-store from training instances and implements a retrieval mechanism during the process of input, training and inference, thus equipping the model with the ability to retrieve related contexts from the training corpus as cues for enhance-ment. Extensive experiments demonstrate that RETROPROMPT can obtain better performance in both few-shot and zero-shot settings. Besides, we further illustrate that our proposed RETROPROMPT can yield better generalization abilities with new datasets. Detailed analysis of memorization indeed reveals RETROPROMPT can reduce the reliance of language models on memorization; thus, improving generalization for downstream tasks.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries of Prompt Learning
  • 3 RETROPROMPT: Retrieval-augmented Prompt Learning
  • 3.1 Dense Retriever
  • 3.2 Retrieval of Neural Demonstration
  • 3.3 Retrieve k NN for Guiding Training
  • 3.4 k NN based probability for Cloze-style Prediction
  • 4 Experiments
  • 4.1 Datasets and Baselines
  • 4.2 Evaluation protocols and details
  • 4.3 Experimental Results
  • 4.4 Model Generalization to New Domains
  • 4.5 Analysis of Memorization
  • 4.6 Ablation Study
  • 5 Related Work
  • 6 Conclusion and Future Work
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — RETROPROMPT Retrieval-Augmented Prompt Learning Framework

    model/method

    RETROPROMPT is a retrieval-augmented prompt tuning framework for pre-trained masked language models (PLMs) designed to decouple task knowledge from parameter memorization. Rather than relying purely on parametric memory or discrete verbalized prompts, RETROPROMPT constructs an open-book key-value datastore (K,V)(K, V) from training instances, where the keys are prompt-based continuous hidden representations and the values are verbalized label tokens.

    The framework integrates non-parametric retrieval into prompt learning across three stages:

    1. Input Augmentation via Neural Demonstrations: Contextual embeddings of the top-mm nearest neighbor examples per class retrieved from the datastore are aggregated via attention weights and concatenated directly to the query embedding sequence.
    2. kNNk\text{NN}-Guided Training: During training, a kk-nearest neighbor distribution over the datastore identifies hard (atypical) instances and dynamically scales the cross-entropy training loss.
    3. kNNk\text{NN}-Interpolated Cloze Prediction: At inference time, the model computes the final prediction by linearly interpolating the PLM's masked language modeling probability distribution over label words with the non-parametric nearest neighbor distribution over the datastore.
  2. Knowl 2 — Neural Demonstration Retrieval and Embedding-Level Concatenation

    model/method

    In RETROPROMPT, query inputs are augmented with continuous representations of retrieved training examples rather than textual prompt demonstrations. For a prompt-formatted query qtq_t with [MASK][\text{MASK}] hidden vector hq^th_{\hat{q}_t}, the dense retriever retrieves the top-mm nearest neighbor examples {c1(l),…,cm(l)}\{c_1^{(l)}, \dots, c_m^{(l)}\} from the open-book knowledge store for each target class l∈{1,…,L}l \in \{1, \dots, L\}, where LL is the total number of classes.

    For each retrieved training instance ci(l)c_i^{(l)}, the knowledge store provides its cached [MASK][\text{MASK}] representation hc^i(l)∈Rdh_{\hat{c}_i}^{(l)} \in \mathbb{R}^d and label word embedding e(v(l))e(v^{(l)}), where e(⋅)e(\cdot) is the PLM's token embedding function. For each class ll, the neighbor representations are combined using softmax attention weights αi(l)\alpha_i^{(l)} based on inner product similarity with hq^th_{\hat{q}_t}: αi(l)=exp⁡(hq^t⊤hc^i(l))∑j=1mexp⁡(hq^t⊤hc^j(l))\alpha_i^{(l)} = \frac{\exp(h_{\hat{q}_t}^\top h_{\hat{c}_i}^{(l)})}{\sum_{j=1}^m \exp(h_{\hat{q}_t}^\top h_{\hat{c}_j}^{(l)})}

    The concatenated input representation II fed into the subsequent Transformer layers of the PLM is formulated as: I=e(x^)⊕[∑i=1mαi(1)hc^i(1),e(v(1))]⊕⋯⊕[∑i=1mαi(L)hc^i(L),e(v(L))]I = e(\hat{x}) \oplus \left[ \sum_{i=1}^m \alpha_i^{(1)} h_{\hat{c}_i}^{(1)}, e(v^{(1)}) \right] \oplus \dots \oplus \left[ \sum_{i=1}^m \alpha_i^{(L)} h_{\hat{c}_i}^{(L)}, e(v^{(L)}) \right] where x^=T(x)\hat{x} = T(x) represents the prompt-wrapped query sentence, and ⊕\oplus denotes sequence concatenation. Because continuous vectors are concatenated, this approach avoids context window overflow in multi-class settings and allows end-to-end gradient backpropagation through the retrieval weights.

  3. Knowl 3 — kNN-Guided Training Loss Calibration

    equation

    In RETROPROMPT, training is guided by a non-parametric kk-nearest neighbor classifier over the open-book knowledge store (K,V)(K, V) to prioritize atypical/hard training examples. For a prompt-transformed training instance qtq_t with [MASK][\text{MASK}] representation hq^th_{\hat{q}_t}, the probability assigned by the kNNk\text{NN} search to class yy is computed over its kk-nearest neighbors N\mathcal{N} (excluding qtq_t via leave-one-out): PkNN(y∣qt)=∑(ci,yi)∈N1y=yiexp⁡(d(hq^t,hc^i))∑(ci,yi)∈Nexp⁡(d(hq^t,hc^i))P_{k\text{NN}}(y \mid q_t) = \frac{\sum_{(c_i, y_i) \in \mathcal{N}} \mathbf{1}_{y = y_i} \exp(d(h_{\hat{q}_t}, h_{\hat{c}_i}))}{\sum_{(c_i, y_i) \in \mathcal{N}} \exp(d(h_{\hat{q}_t}, h_{\hat{c}_i}))} where d(⋅,⋅)d(\cdot, \cdot) denotes the inner product similarity metric, and 1y=yi\mathbf{1}_{y = y_i} is the indicator function.

    Let pkNN=PkNN(ygold∣qt)p_{k\text{NN}} = P_{k\text{NN}}(y_{\text{gold}} \mid q_t) be the kNNk\text{NN} probability assigned to the ground-truth class ygoldy_{\text{gold}}. The modulating factor F(pkNN)F(p_{k\text{NN}}) and the adjusted training loss L\mathcal{L} are defined as: F(pkNN)=−log⁡(pkNN)F(p_{k\text{NN}}) = -\log(p_{k\text{NN}}) L=(1+βF(pkNN))LCE\mathcal{L} = (1 + \beta F(p_{k\text{NN}})) \mathcal{L}_{\text{CE}} where LCE\mathcal{L}_{\text{CE}} is the standard cross-entropy loss and β≥0\beta \ge 0 is a scalar hyperparameter governing the weight of the difficulty modulation. Harder examples where kNNk\text{NN} fails receive a larger loss multiplier.

  4. Knowl 4 — Cloze-Style Inference with kNN Probability Interpolation

    equation

    During inference in RETROPROMPT, prediction combines the parametric output of the masked language model (PLM) with a non-parametric nearest neighbor distribution over the prompt knowledge store (K,V)(K, V).

    For a query qtq_t wrapped by prompt template T(qt)T(q_t), let PM([MASK]=v∣T(qt))P_M([\text{MASK}] = v \mid T(q_t)) denote the MLM head probability of predicting token vv at the [MASK][\text{MASK}] position, and let g(⋅)g(\cdot) map the probability of label words in vocabulary subset Vy\mathcal{V}_y to class y∈Yy \in \mathcal{Y}. Let PkNN(y∣qt)P_{k\text{NN}}(y \mid q_t) be the non-parametric probability computed over the top-kk nearest neighbors in the datastore. The final class probability distribution P(y∣qt)P(y \mid q_t) is computed by linear interpolation: P(y∣qt)=λPkNN(y∣qt)+(1−λ)g(PM([MASK]=v∣T(qt))∣v∈Vy)P(y \mid q_t) = \lambda P_{k\text{NN}}(y \mid q_t) + (1 - \lambda) g\left( P_M([\text{MASK}] = v \mid T(q_t)) \mid v \in \mathcal{V}_y \right) where λ∈[0,1]\lambda \in [0, 1] is an interpolation hyperparameter balancing non-parametric retrieval memory and parametric prompt prediction.

  5. Knowl 5 — Open-Book Knowledge-Store Construction and Asynchronous Refresh

    model/method

    The open-book knowledge store (K,V)(K, V) in RETROPROMPT indexes training examples using dense contextual embeddings derived from prompt templates. For each training sample (ci,yi)∈C(c_i, y_i) \in \mathcal{C}, the input is transformed by template function T(ci)T(c_i) into a cloze sequence c^i\hat{c}_i. The datastore is populated as: (K,V)={(hc^i,vi)∣(ci,yi)∈C}(K, V) = \{(h_{\hat{c}_i}, v_i) \mid (c_i, y_i) \in \mathcal{C}\} where hc^i∈Rdh_{\hat{c}_i} \in \mathbb{R}^d is the final hidden state vector at the [MASK][\text{MASK}] token position extracted from the PLM encoder, and vi=f(yi)v_i = f(y_i) is the verbalized label word for label yiy_i. Nearest neighbor queries are executed over key matrix D∈R∣C∣×dD \in \mathbb{R}^{|\mathcal{C}| \times d} using Maximum Inner Product Search (MIPS) via FAISS.

    To account for representation drift as PLM parameters update during training, the index is refreshed asynchronously every jj training epochs (typically j=1j=1) by re-embedding all training instances with the current model weights and updating the FAISS index.

    In zero-shot settings without labeled data, the datastore is constructed using an unlabeled training corpus where pseudo-labels are assigned by a zero-shot base PLM before building (K,V)(K, V), and predictions are made via retrieval and base prompt inference without parameter updates.

  6. Knowl 6 — Self-Influence Memorization Measurement for Training Instances

    definition

    To quantify the degree of memorization exhibited by language models for individual training instances, RETROPROMPT employs a sample deletion influence metric based on self-influence. For a training instance z=(x,y)z = (x, y), memorization score Sdelete(z)S_{\text{delete}}(z) is defined as the rate of change in the predicted probability P(y∣x;θ^ξ,−z)P(y \mid x; \hat{\theta}_{\xi, -z}) as the sample zz is down-weighted by a factor ξ→0\xi \to 0: Sdelete(z)=def−dP(y∣x;θ^ξ,−z)dξ∣ξ=0=−∇θP(y∣x;θ^)⊤Hθ^−1∇θL(z,θ^)S_{\text{delete}}(z) \stackrel{\text{def}}{=} -\left. \frac{d P(y \mid x; \hat{\theta}_{\xi, -z})}{d\xi} \right|_{\xi=0} = -\nabla_\theta P(y \mid x; \hat{\theta})^\top H_{\hat{\theta}}^{-1} \nabla_\theta \mathcal{L}(z, \hat{\theta}) where θ^\hat{\theta} represents the model parameters trained on the full dataset, θ^ξ,−z\hat{\theta}_{\xi, -z} represents parameters trained with instance zz weighted by (1−ξ)(1-\xi), L(z,θ^)\mathcal{L}(z, \hat{\theta}) is the loss on sample zz, and Hθ^=1n∑i=1n∇θ2L(zi,θ^)H_{\hat{\theta}} = \frac{1}{n} \sum_{i=1}^n \nabla_\theta^2 \mathcal{L}(z_i, \hat{\theta}) is the empirical Hessian matrix evaluated over all nn training instances. A higher value of Sdelete(z)S_{\text{delete}}(z) indicates that the model's prediction for instance zz depends heavily on rote memorization of zz during training.

  7. Knowl 7 — Few-Shot and Zero-Shot Performance Across Nine NLU Benchmarks

    data/table

    RETROPROMPT was evaluated using RoBERTa-large on 9 natural language understanding datasets across single-sentence classification (SST-2, MR, CR), sentence-pair classification (MNLI, QNLI, QQP), and information extraction (FewNERD, SemEval-2010 Task 8, TACRED) in 16-shot, 4-shot, and zero-shot settings. The baseline models include standard fine-tuning (FT), LM-BFF (manual and discrete demonstration), KnowPrompt (KnPr), LOTClass, and Knowledgeable Prompt-tuning (KPT).

    Shot Model SST-2 MR CR MNLI QNLI QQP FewNERD SemEval TACRED
    (acc) (acc) (acc) (acc) (acc) (F1) (acc) (acc) (F1)
    16 FT 81.4 76.9 75.8 45.8 60.2 60.7 52.7 66.1 25.8
    LM-BFF (man) 91.6 87.0 90.3 64.3 64.6 65.4 — — —
    LM-BFF (D-demo) 91.8 86.6 90.2 64.8 69.2 68.2 — — —
    KnowPrompt — — — — — — 65.3 80.9 33.2
    KPT 90.3 86.8 88.8 61.4 61.5 71.6 65.9 78.8 32.8
    RETROPROMPT 93.9 88.0 91.9 71.1 71.6 74.0 67.3 81.5 40.7
    4 FT 60.2 57.6 66.4 35.0 54.2 52.8 32.7 38.8 14.7
    LM-BFF (man) 90.7 85.2 89.9 51.0 61.1 48.0 — — —
    LM-BFF (D-demo) 90.2 85.5 89.7 56.1 61.7 63.2 — — —
    KnowPrompt — — — — — — 52.5 58.4 28.8
    KPT 88.2 83.4 87.2 53.7 59.2 54.9 58.8 57.2 27.5
    RETROPROMPT 91.5 87.4 91.4 57.6 62.2 66.1 60.9 59.2 32.1
    0 LOTClass 71.8 81.7 50.1 50.4 36.5 55.9 11.5 9.8 2.5
    FT 49.1 50.0 49.8 34.4 49.5 31.6 10.0 6.2 0.5
    LM-BFF (man) 83.5 80.3 78.4 49.7 50.5 49.7 — — —
    LM-BFF (D-demo) 82.9 80.7 81.4 52.2 53.5 44.0 — — —
    KnowPrompt — — — — — — 15.9 10.3 2.3
    KPT 78.4 81.9 71.4 37.1 55.3 47.5 24.6 11.6 0.8
    RETROPROMPT 86.8 83.5 79.7 53.7 56.2 56.7 41.3 12.2 2.8

    Across all settings (16-shot, 4-shot, and zero-shot), RETROPROMPT achieves the highest average score (75.6 in 16-shot, 67.6 in 4-shot, and 52.5 in zero-shot) while outperforming both external knowledge-augmented prompt learning (KPT) and discrete demonstration prompt baselines.

  8. Knowl 8 — Cross-Domain Generalization of RETROPROMPT

    data/table

    To evaluate domain transfer ability without overfitting to source distribution artifacts, models trained on source datasets in the 16-shot setting were directly evaluated on target domain datasets without retraining or fine-tuning.

    Model Source Domain Target Domain
    16-shot MR SST-2 CR
    FT 76.9 71.4 64.7
    LM-BFF (man) 87.0 88.9 86.9
    LM-BFF (D-demo) 86.6 89.3 87.5
    KPT 86.8 86.8 86.7
    RETROPROMPT 88.0 91.4 88.8
    16-shot QQP MRPC RTE
    FT 60.7 43.7 48.0
    LM-BFF (man) 65.4 20.9 65.5
    LM-BFF (D-demo) 68.2 38.8 66.2
    KPT 71.6 42.3 65.8
    RETROPROMPT 74.0 49.4 67.3

    RETROPROMPT outperforms fine-tuning and prompt-tuning baselines on all transfer pairs: from MR to SST-2 (91.4% vs. 89.3% best baseline) and CR (88.8% vs. 87.5%), and from QQP to MRPC (49.4% vs. 42.3%) and RTE (67.3% vs. 66.2%).

  9. Knowl 9 — Reduction of Memorization Dependency on Atypical Samples

    empirical result

    Empirical analysis on SST-2 demonstrates that pre-trained language models assign the highest memorization scores (SdeleteS_{\text{delete}}) to atypical training instances (e.g., negative sentences containing a high percentage of positive sub-phrases).

    Memory Group Negative Instances (% Positive Phrases) Positive Instances (% Positive Phrases)
    FT LM-BFF RETROPROMPT FT LM-BFF RETROPROMPT
    Top-10% 34.29 32.78 30.23 68.75 69.71 75.67
    All Instances 23.40 86.39
    Bottom-10% 17.63 16.25 14.42 95.92 95.08 94.53

    For negative instances (corpus average 23.40% positive phrases), the Top-10% highest-memorization instances have 30.23% to 34.29% positive phrases, indicating atypicality. Conversely, the Bottom-10% memorized instances have only 14.42% to 17.63% positive phrases (typical).

    Across all training samples in SST-2, the mean memorization score drops substantially from standard fine-tuning (4.597) to LM-BFF (0.121) to RETROPROMPT (0.032). This confirms that decoupling knowledge into an external retrieval store reduces the model's parametric reliance on rote memorization.

  10. Knowl 10 — Component Ablations and Key Representation Comparison

    empirical result

    Ablation studies across 16-shot classification and extraction tasks evaluate the individual contributions of RETROPROMPT's components:

    • w/o kNN-test: removes kNNk\text{NN} probability interpolation during Cloze-style test inference.
    • w/o kNN-train: removes kNNk\text{NN}-guided loss modulating factor F(pkNN)F(p_{k\text{NN}}).
    • w/o N-demo: removes neural demonstration embeddings from input sequences.
    • w/o refresh: disables asynchronous re-indexing of knowledge-store embeddings across training epochs.
    Ablation Variant SST-2 CR MNLI QQP TACRED
    RETROPROMPT (Full) 93.9 91.9 71.1 74.0 40.7
    w/o kNN-test 93.2 91.2 70.4 73.0 38.2
    w/o kNN-train 92.0 90.2 68.8 71.3 36.5
    w/o N-demo 92.4 91.0 70.1 72.7 37.9
    w/o refresh 93.5 91.5 70.7 73.6 39.9

    The ablation demonstrates that kNN-train and N-demo provide the largest performance boosts in the few-shot regime.

    Additionally, comparing key representation schemes and similarity metrics on 16-shot CR and TACRED shows that using prompt [MASK][\text{MASK}] embeddings with inner product similarity achieves higher accuracy/F1 (91.9 / 40.7) compared to [CLS][\text{CLS}] embeddings with inner product similarity (89.0 / 37.2), prompt embeddings with BM25 (89.5 / 38.8), or [CLS][\text{CLS}] embeddings with BM25 (88.7 / 36.1).

Coverage note — Qualitative instance case studies from Table 6 were omitted in favor of quantitative memorization analysis and ablation tables.

References

  1. 1.Uri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig. Neuro-symbolic language modeling with automaton-augmented retrieval, 2022.
  2. 2.Gianluca Bontempi, Hugues Bersini, and Mauro Birattari. The local paradigm for modeling and control: from neuro-fuzzy to lazy learning. Fuzzy sets and systems, 121(1):59–72, 2001.
  3. 3.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. Improving language models by retrieving from trillions of tokens. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 2206–2240. PMLR, 2022.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Proceedings of NeurIPS 2020, 2020.
  5. 5.Xiang Chen, Ningyu Zhang, Lei Li, Xin Xie, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. Lightner: A lightweight generative framework with prompt-guided attention for low-resource NER. CoRR, abs/2109.00720, 2021.
  6. 6.Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction. CoRR, abs/2104.07650, 2021.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019.
  8. 8.Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. Few-nerd: A few-shot named entity recognition dataset. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3198–3213. Association for Computational Linguistics, 2021.
  9. 9.Aparna Elangovan, Jiayuan He, and Karin Verspoor. Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1325–1335, Online, April 2021. Association for Computational Linguistics.
  10. 10.Vitaly Feldman. Does learning require memorization? a short tale about a long tail. In Konstantin Makarychev, Yury Makarychev, Madhur Tulsiani, Gautam Kamath, and Julia Chuzhoy, editors, Proccedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, STOC 2020, Chicago, IL, USA, June 22-26, 2020, pages 954–959. ACM, 2020.
  11. 11.Vitaly Feldman and Chiyuan Zhang. What neural networks memorize and why: Discovering the long tail via influence estimation. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  12. 12.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In Proceedings of ACL, 2021.
  13. 13.Edouard Grave, Moustapha Cissé, and Armand Joulin. Unbounded cache model for online language modeling with open vocabulary. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6042–6052, 2017.
  14. 14.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909, 2020.
  15. 15.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. PTR: prompt tuning with rules for text classification. CoRR, abs/2105.11259, 2021.
  16. 16.Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick. Efficient nearest neighbor language models. In Proc. of EMNLP, 2021.
  17. 17.Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals. In Proceedings of SemEval, pages 33–38, 2010.
  18. 18.I Hsu, Kuan-Hao Huang, Elizabeth Boschee, Scott Miller, Prem Natarajan, Kai-Wei Chang, Nanyun Peng, et al. Event extraction as natural language generation. arXiv preprint arXiv:2108.12724, 2021.
  19. 19.Minqing Hu and Bing Liu. Mining and summarizing customer reviews. In Won Kim, Ron Kohavi, Johannes Gehrke, and William DuMouchel, editors, Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Seattle, Washington, USA, August 22-25, 2004, pages 168–177. ACM, 2004.
  20. 20.Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Juanzi Li, and Maosong Sun. Knowledgeable prompt-tuning: Incorporating knowledge into prompt verbalizer for text classification. CoRR, abs/2108.02035, 2021.
  21. 21.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with gpus. IEEE Trans. Big Data, 7(3):535–547, 2021.
  22. 22.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. Spanbert: Improving pre-training by representing and predicting spans. Trans. Assoc. Comput. Linguistics, 8:64–77, 2020.
  23. 23.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Association for Computational Linguistics, 2020.
  24. 24.Nora Kassner and Hinrich Schütze. Bert-knn: Adding a knn search component to pretrained language models for better QA. In Findings of EMNLP, 2020.
  25. 25.Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Nearest neighbor machine translation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  26. 26.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: Nearest neighbor language models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020.
  27. 27.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. In International Conference on Machine Learning, 2017.
  28. 28.Sawan Kumar and Partha Talukdar. Reordering examples helps during priming-based few-shot learning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4507–4518, Online, 2021. Association for Computational Linguistics.
  29. 29.Dong-Ho Lee, Mahak Agarwal, Akshen Kadakia, Jay Pujara, and Xiang Ren. Good examples make A faster learner: Simple demonstration-based learning for low-resource NER. CoRR, abs/2110.08454, 2021.
  30. 30.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  31. 31.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL 2020, 2020.
  32. 32.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  33. 33.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of ACL/IJCNLP 2021, 2021.
  34. 34.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In IEEE Trans. Pattern Anal. Mach. Intell., volume 42, pages 318–327, 2020.
  35. 35.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? CoRR, abs/2101.06804, 2021.
  36. 36.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586, 2021.
  37. 37.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. GPT understands, too. CoRR, abs/2103.10385, 2021.
  38. 38.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019.
  39. 39.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. CoRR, abs/2104.08786, 2021.
  40. 40.Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Qi Zhang, and Xuanjing Huang. Template-free prompt tuning for few-shot NER. CoRR, abs/2109.13532, 2021.
  41. 41.Yu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong, Heng Ji, Chao Zhang, and Jiawei Han. Text classification using label names only: A language model self-training approach. In Proceedings of EMNLP, 2020.
  42. 42.Yuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun, Tianwei Zhang, Fei Wu, and Jiwei Li. GNN-LM: language modeling based on global contexts via GNN. CoRR, abs/2110.08743, 2021.
  43. 43.Bo Pang and Lillian Lee. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Kevin Knight, Hwee Tou Ng, and Kemal Oflazer, editors, ACL 2005, 43rd Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, 25-30 June 2005, University of Michigan, USA, pages 115–124. The Association for Computer Linguistics, 2005.
  44. 44.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 8024–8035, 2019.
  45. 45.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. OpenAI, 2018.
  46. 46.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100, 000+ questions for machine comprehension of text. In Jian Su, Xavier Carreras, and Kevin Duh, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392. The Association for Computational Linguistics, 2016.
  47. 47.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3980–3990. Association for Computational Linguistics, 2019.
  48. 48.Stephen E. Robertson and Hugo Zaragoza. The probabilistic relevance framework: BM25 and beyond. Found. Trends Inf. Retr., 3(4):333–389, 2009.
  49. 49.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning to retrieve prompts for in-context learning. CoRR, abs/2112.08633, 2021.
  50. 50.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M. Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. Multitask prompted training enables zero-shot task generalization. CoRR, abs/2110.08207, 2021.
  51. 51.Timo Schick, Helmut Schmid, and Hinrich Schütze. Automatically identifying words that can serve as labels for few-shot text classification. In Proceedings of COLING, December 2020.
  52. 52.Timo Schick and Hinrich Schütze. Few-shot text generation with pattern-exploiting training. arXiv preprint arXiv:2012.11926, 2020.
  53. 53.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. Auto-prompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of EMNLP 2020, 2020.
  54. 54.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1631–1642. ACL, 2013.
  55. 55.Zhixing Tan, Xiangwen Zhang, Shuo Wang, and Yang Liu. MSP: multi-stage prompting for making pre-trained language models better translators. CoRR, abs/2110.06609, 2021.
  56. 56.Michael Tänzer, Sebastian Ruder, and Marek Rei. Memorisation versus generalisation in pre-trained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7564–7578, 2022.
  57. 57.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. CoRR, abs/2109.01652, 2021.
  58. 58.Adina Williams, Nikita Nangia, and Samuel R. Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Marilyn A. Walker, Heng Ji, and Amanda Stent, editors, Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics, 2018.
  59. 59.Sen Yang, Yunchen Zhang, Leyang Cui, and Yue Zhang. Do prompts solve NLP tasks using natural language? CoRR, abs/2203.00902, 2022.
  60. 60.Hongbin Ye, Ningyu Zhang, Zhen Bi, Shumin Deng, Chuanqi Tan, Hui Chen, Fei Huang, and Huajun Chen. Learning to ask for data-efficient event argument extraction. arXiv preprint arXiv:2110.00479, 2021.
  61. 61.Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning. Position-aware attention and supervised data improve slot filling. In Proceedings of EMNLP 2017, 2017.
  62. 62.Xiaosen Zheng and Jing Jiang. An empirical study of memorization in NLP. CoRR, abs/2203.12171, 2022.

Citation

MLA
Chen, X., et al. “Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 23908–22, https://proceedings.neurips.cc/paper_files/paper/2022/file/97011c648eda678424f9292dadeae72e-Paper-Conference.pdf.
APA
Chen, X., Li, L., Zhang, N., Liang, X., Deng, S., Tan, C., Huang, F., Si, L., & Chen, H. (2022). Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning. Advances in Neural Information Processing Systems, 35, 23908–23922. https://proceedings.neurips.cc/paper_files/paper/2022/file/97011c648eda678424f9292dadeae72e-Paper-Conference.pdf
Chicago
Chen, X., L. Li, N. Zhang, et al. 2022. “Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning”. Advances in Neural Information Processing Systems 35: 23908–22. https://proceedings.neurips.cc/paper_files/paper/2022/file/97011c648eda678424f9292dadeae72e-Paper-Conference.pdf.
Harvard
Chen, X. et al. (2022) “Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 23908–23922. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/97011c648eda678424f9292dadeae72e-Paper-Conference.pdf.
Vancouver
1. Chen X, Li L, Zhang N, Liang X, Deng S, Tan C, Huang F, Si L, Chen H (2022) Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 23908–23922

BibTeX

@inproceedings{chen2022decoupling,
  title = {Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning},
  author = {Chen, Xiang and Li, Lei and Zhang, Ningyu and Liang, Xiaozhuan and Deng, Shumin and Tan, Chuanqi and Huang, Fei and Si, Luo and Chen, Huajun},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {23908-23922},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/97011c648eda678424f9292dadeae72e-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors