Memory-assisted prompt editing to improve GPT-3 after deployment

Aman MadaanNiket TandonPeter ClarkYiming Yang

article2022EMNLP183 citations

Proposes an interactive framework that pairs deployed large language models with an external memory of user feedback to dynamically update prompts and correct instruction misunderstandings without costly retraining.

Listen

Large language models such as GPT-3 often misinterpret user instructions when questions are phrased ambiguously, use varied dialects, or rely on novel wording. While traditional model improvement relies on retraining or fine-tuning, doing so after deployment is prohibitively expensive and technically demanding. Consequently, deployed systems frequently repeat the same interpretive mistakes without any practical mechanism to learn from user corrections.

The article demonstrates and evaluates an interactive framework called MEM-PROMPT, which enables a deployed language model to correct instruction misunderstandings dynamically through prompt editing and feedback memory, entirely avoiding model retraining.

The authors tested this approach using GPT-3 on ten tasks divided into two categories: five lexical question-answering tasks (synonyms, antonyms, homonyms, definitions, and sentence generation) and five word-scrambling tasks. The architecture prompts the model to articulate its interpretation of user intent alongside its answer. If the interpretation is flawed, user-provided feedback is stored in an external key-value memory. For incoming queries, a similarity retriever identifies past relevant corrections and automatically prepends them to the prompt. Evaluated over 300 data points per task, MEM-PROMPT was benchmarked against standard few-shot GPT-3 (NO-MEM) and a naive memory approach (GROW-PROMPT) that simply appends raw past feedback up to context limits.

The evaluation yielded several key results. First, MEM-PROMPT raised overall accuracy on lexical tasks from 37% to 98%, more than doubling the performance of standard GPT-3. Second, it outperformed the naive GROW-PROMPT baseline (which achieved 80% accuracy) while avoiding context-window exhaustion and reducing prompt costs by approximately two-thirds. Third, the retrieved corrective feedback proved effective 97% of the time, failing to help in only 3% of cases. Fourth, the system showed substantial accuracy gains in multilingual settings where queries were posed in transcribed Hindi and Punjabi with English feedback. Finally, accuracy gains were smaller on word-scrambling tasks (rising from 77% to 90%) because those instructions inherently had lower semantic ambiguity.

These findings indicate that deployed artificial intelligence systems can adaptively correct operational errors without expensive retraining cycles. By having the model state its understanding of an instruction, end users can provide helpful guidance on intent even without knowing the correct final answer, substantially lowering the technical barrier for user feedback. This dynamic memory mechanism allows systems to continuously improve from real-world usage while keeping operational inference costs manageable.

Organizations deploying large language models should consider integrating memory-assisted prompt-editing architectures into production workflows where prompt ambiguity and varied user phrasing cause repetitive errors. Future work should focus on developing more sophisticated gating mechanisms to filter out irrelevant feedback, optimizing retrieval algorithms (such as dense vector indices) for larger memory tables, and validating the framework in complex conversational settings beyond structured lexical benchmarks.

The findings are supported by consistent proof-of-concept experiments, but confidence is bounded by specific limitations: the study used synthetic and template-based user feedback, evaluated relatively narrow lexical and word-scrambling tasks, and used a limited sample size of 300 evaluations per benchmark. Stakeholders should conduct pilot testing in realistic, diverse conversational environments before full-scale deployment.

Abstract

Large LMs such as GPT-3, while powerful, are not immune to mistakes, but are prohibitively costly to retrain. One failure mode is misinterpreting a user’s instruction (e.g., GPT-3 interpreting "What word is similar to good?" to mean a homonym, while the user intended a synonym). Our goal is to allow users to correct such errors directly through interaction – without retraining. Our approach pairs GPT-3 with a growing memory of cases where the model misunderstood the user’s intent and was provided with feedback, clarifying the instruction. Given a new query, our memory-enhanced GPT-3 uses feedback from similar, prior queries to enrich the prompt. Through simple proof-of-concept experiments, we show how a (simulated) user can interactively teach a deployed GPT-3, doubling its accuracy on basic lexical tasks (e.g., generate a synonym) where users query in different, novel (often misunderstood) ways. In such scenarios, memory helps avoid repeating similar past mistakes. Our simple idea is a first step towards strengthening deployed models, potentially broadening their utility.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Approach
  • 3.1 Memory enhanced gpt-3 architecture
  • 3.2 Tasks
  • 3.3 A Proof of Concept Implementation
  • 4 Experiments
  • 5 Conclusion
  • References
  • A Querying gpt-3-175b using OpenAI API
  • B Prompt
  • C Datasets for lexical question-answering tasks
  • C.1 Templates
  • C.2 Sample questions
  • D Finding similar questions in low-resource settings
  • E Sample results

Knowls

  1. Knowl 1 — MEM-PROMPT Framework for Interactive Prompt Editing

    model/method

    The MEM-PROMPT architecture allows a deployed, fixed large language model (GPT-3) to correct instruction misunderstandings post-deployment through natural language user feedback without retraining. The system consists of four primary components:

    1. Memory M\mathcal{M}: A dynamic key-value table storing past error instances where the key is a misunderstood user query xix_i and the value is corrective user feedback fbifb_i that clarifies the user's intent. M\mathcal{M} supports read, write, and lookup operations.
    2. Lookup Retriever Ω(x,M)\Omega(x, \mathcal{M}): A retriever that searches M\mathcal{M} for keys matching the incoming query xx, relying on the intent exchangeability principle: if two queries share similar error modes (xi∼xjx_i \sim x_j), their clarifying feedbacks are interchangeable (fbi∼fbjfb_i \sim fb_j).
    3. Combiner C(x,Ω(x,M))\mathcal{C}(x, \Omega(x, \mathcal{M})): A gating and formatting function that concatenates the input query xx with the retrieved feedback Ω(x,M)\Omega(x, \mathcal{M}). If no relevant feedback is retrieved, the query remains unaugmented.
    4. Prompter P(p,C)\mathcal{P}(p, \mathcal{C}): Appends the combined query and feedback to a base few-shot prompt pp, conditioning GPT-3 to produce both its understood task intent uu and the answer yy.

    When a user observes that the model generated an incorrect intent understanding uu, new feedback fbnewfb_{\text{new}} is written into M\mathcal{M}, allowing the model to adapt dynamically over continuous user interactions.

  2. Knowl 2 — Intent Verbalization for Non-Expert User Feedback

    model/method

    To enable feedback from non-expert users who may not know the correct ground-truth answer y∗y^* for a task (e.g., in complex translation or specialized definition tasks), the model is prompted via few-shot demonstrations to generate both an explicit verbalization of its understood task intent uu alongside the task answer yy:

    Model output=[u]  [y]\text{Model output} = [u] \; [y]

    For example, given a query x="What is akin to fast?"x = \text{"What is akin to fast?"}, an unclarified model might generate u="The opposite of fast is:"u = \text{"The opposite of fast is:"} and y="slow"y = \text{"slow"}.

    By inspecting uu, an end-user detects that the model misinterpreted the intent as requesting an antonym rather than a synonym. The user can then provide corrective natural language feedback fbfb (e.g., fb="clarification: when I ask for like, I want a synonym."fb = \text{"clarification: when I ask for like, I want a synonym."}) critiquing solely the intent interpretation uu, without needing domain expertise to supply the ground-truth output y∗y^*.

  3. Knowl 3 — MEM-PROMPT Retrieval and Query Execution Procedure

    algorithm
    Input: Input query xx, static few-shot prompt pp, dynamic memory table M={(xk,fbk)}k=1∣M∣\mathcal{M} = \{(x_k, fb_k)\}_{k=1}^{|\mathcal{M}|}, similarity threshold τ=0.9\tau = 0.9, feedback probability P(fb)P(fb)
    Output: Generated intent uu, answer yy, updated memory M\mathcal{M}
    1. Feedback Retrieval:
       a. If M\mathcal{M} is not empty, compute similarity s(x,xk)s(x, x_k) for all (xk,fbk)∈M(x_k, fb_k) \in \mathcal{M}.
          - In standard semantic mode: s(x,xk)s(x, x_k) is the cosine similarity of Sentence-BERT embeddings.
          - In low-resource mode: s(x,xk)s(x, x_k) is based on Levenshtein distance over surface tokens.
       b. Identify the closest match xm=argmax⁡xks(x,xk)x_m = \operatorname*{argmax}_{x_k} s(x, x_k).
       c. If s(x,xm)≥τs(x, x_m) \ge \tau, set retrieved feedback fb=fbmfb = fb_m; otherwise set fb=∅fb = \emptyset.
    2. Prompt Construction and Model Inference:
       a. If fb≠∅fb \neq \emptyset, format augmented query xaug=x+" | clarification: "+fbx_{\text{aug}} = x + \text{" | clarification: "} + fb.
       b. Else, format xaug=xx_{\text{aug}} = x.
       c. Feed p#xaugp \mathbin{\#} x_{\text{aug}} to GPT-3 (davinci-175B).
       d. Parse continuation into intent uu and task output yy.
    3. Dynamic Memory Update:
       a. If intent uu is judged incorrect by user, with probability P(fb)P(fb) obtain corrective feedback fbnewfb_{\text{new}}.
       b. If fbnewfb_{\text{new}} is provided, write key-value pair (x,fbnew)(x, fb_{\text{new}}) into M\mathcal{M}.
    return u,y,Mu, y, \mathcal{M}
  4. Knowl 4 — In-Context Prompt Structure for Feedback Conditioning

    model/method

    To train GPT-3 (davinci-175B) in-context to follow corrective feedback at inference time without parameter updates, the static base prompt pp is constructed with kk few-shot exemplar tuples containing a mixture of unaugmented demonstrations (x→u,y)(x \to u, y) and feedback-augmented demonstrations (x,fb→u,y)(x, fb \to u, y) separated by a delimiter token #:

    • Standard example: x # u y END #
    • Feedback-augmented example: x | clarification: fb # u y END #

    For example, an exemplar demonstrating feedback reaction is: what is like < provident > ? | clarification: when I ask for like, I want a synonym. # the synonym for provident is prudent END #

    When a new query is suffixed with retrieved feedback (xi∣clarification: fbix_i \mid \text{clarification: } fb_i), GPT-3 matches the in-context demonstration pattern, causing the model to adapt its generated understanding uiu_i to align with fbifb_i and output the correct response yiy_i. Generation parameters are set to: temperature=0.7, max_tokens=64, top_p=1, frequency_penalty=0, and presence_penalty=0.

  5. Knowl 5 — Baseline Comparison and Accuracy on Lexical QA Tasks

    data/table

    Performance evaluated over 300 sequential test instances across five lexical QA tasks (synonym syn, antonym ant, homonym hom, sentence usage generation sent, and definition defn) comparing MEM-PROMPT against standard few-shot GPT-3 (NO-MEM) and a non-selective memory baseline (GROW-PROMPT) that appends all recent feedback pairs into the prompt up to the 2048 token limit:

    Model syn ant hom sent defn All
    NO-MEM 0.58 0.43 0.13 0.30 0.39 0.37
    GROW-PROMPT 0.71 0.87 0.75 0.92 0.76 0.80
    MEM-PROMPT 0.99 0.98 0.98 0.98 0.96 0.98

    MEM-PROMPT achieves 0.980.98 overall accuracy across 300 steps, more than doubling the 0.370.37 accuracy of NO-MEM. GROW-PROMPT reaches 0.800.80 accuracy but is approximately 3×3\times more expensive in prompt token consumption and cannot scale beyond the 2048-token context window.

  6. Knowl 6 — Performance on Word Scrambling and Character Manipulation Tasks

    data/table

    Performance evaluated over 300 test instances across five word scrambling and character manipulation tasks: first-and-last character anagrams (anag1), first-and-last-two character anagrams (anag2), letter cycling (cyc), random punctuation insertion (rand), and word reversal (rev):

    Model anag1 anag2 cyc rand rev All
    NO-MEM 0.81 0.47 0.95 0.98 0.62 0.77
    GROW-PROMPT 0.86 0.89 0.93 0.96 0.90 0.91
    MEM-PROMPT 0.81 0.83 0.98 0.95 0.93 0.90

    Both memory-assisted approaches substantially improve over the unaugmented NO-MEM baseline (0.770.77). The accuracy improvement for MEM-PROMPT (0.900.90) is less drastic than in lexical QA (0.980.98) because word scrambling task instructions are less ambiguous, resulting in fewer initial misinterpretations of user intent by the base model.

  7. Knowl 7 — Lexical QA Benchmark Task Construction and Ambiguity Design

    experimental setup

    The lexical QA experimental evaluation is constructed using 15 manual task templates with three phrasing variations per task across five distinct linguistic categories:

    1. Synonyms (syn) and Antonyms (ant): Sourced from the lexical contrast dataset of Nguyen et al. (2016).
    2. Homonyms (hom): Defined as pairs with identical pronunciations but differing spellings (e.g., ring vs. wring), generated via the homz dictionary lookup.
    3. Definitions (defn): Extracted from The Online Plain Text English Dictionary.
    4. Sentence usage generation (sent): Extracted from the CommonGen dataset.

    Questions are deliberately constructed using ambiguous phrasing (e.g., for synonyms: "What is like <word>?", "What has a similar sense?", "What is akin to <word>?") so that the base language model frequently confuses similar intent categories (e.g., confusing "what sounds like" with "what is like") unless clarified by user feedback fbfb.

  8. Knowl 8 — Cross-Lingual Generalization and Surface-Form Retrieval in OOV Settings

    empirical result

    MEM-PROMPT demonstrates effective out-of-vocabulary (OOV) and code-mixed generalization when queries are formulated in transliterated Hindi (e.g., <tabulate> ka ulta kya hai? for antonyms) or Punjabi (e.g., <edit> de ult ki hunda ae?), accompanied by clarifying feedback provided in English.

    In low-resource or transliterated settings where pre-trained dense sentence encoders fail, MEM-PROMPT implements memory lookup Ω(x,M)\Omega(x, \mathcal{M}) via Levenshtein distance surface-form matching. Under this cross-lingual setting, standard GPT-3 (NO-MEM) fails predictably due to unfamiliar phrasing, whereas MEM-PROMPT leverages past English feedback indexed under similar surface strings to correct intent understanding and rapidly achieve high accuracy within 300 steps.

  9. Knowl 9 — Memory Feedback Efficacy and Retrieval Precision

    empirical result

    Quantitative analysis of the MEM-PROMPT feedback and retrieval mechanism reveals three key operational properties:

    1. Feedback Retrieval Efficacy: Retrieved feedback from memory M\mathcal{M} successfully corrects the model's behavior 97%97\% of the time; in only ≈3%\approx 3\% of cases does retrieved feedback fail to produce a positive effect.
    2. Feedback Persistence Rate (P(fb)P(fb)): Continuous feedback (P(fb)=1.0P(fb) = 1.0) accelerates the rate of accuracy improvement across 300 steps compared to intermittent feedback (P(fb)=0.5P(fb) = 0.5), although intermittent feedback still significantly outperforms memory-free baselines.
    3. Intent-to-Output Correlation: Across all lexical tasks, there exists a near-perfect correlation between the correctness of the generated intent verbalization uu and the final answer yy: whenever GPT-3 generates the correct understanding uu, the task output yy is almost always correct.

Coverage note — None omitted; all primary architectural components, algorithmic mechanisms, empirical evaluation results, benchmark task specifications, and ablations have been captured.

References

  1. 1.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, and Others. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  2. 2.Randall Davis. 1977. Interactive transfer of expertise: Acquisition of new inference rules. Artif. Intell., 12:121–157.
  3. 3.Kelvin Guu, Kenton Lee, Z. Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. ArXiv, abs/2002.08909.
  4. 4.Ben Hixon, Peter Clark, and Hannaneh Hajishirzi. 2015. Learning knowledge graphs for question answering through conversational dialog. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 851–861, Denver, Colorado. Association for Computational Linguistics.
  5. 5.Larry L. Jacoby and Christopher N. Wahlheim. 2013. On the importance of looking back: The role of recursive remindings in recency judgments and cued recall. Memory & Cognition, 41:625–637.
  6. 6.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017. Billion-scale similarity search with gpus. arXiv preprint arXiv:1702.08734.
  7. 7.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  8. 8.Teven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636, Online. Association for Computational Linguistics.
  9. 9.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  10. 10.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  11. 11.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021a. What makes good in-context examples for gpt-3? ArXiv, abs/2101.06804.
  12. 12.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021b. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ArXiv.
  13. 13.Gary Marcus. Experiments testing gpt-3’s ability at commonsense reasoning: results. [online]. 2021.
  14. 14.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Natural instructions: Benchmarking generalization to new tasks from natural language instructions. ArXiv, abs/2104.08773.
  15. 15.Kim Anh Nguyen, Sabine Schulte im Walde, and Ngoc Thang Vu. 2016. Integrating distributional lexical contrast into word embeddings for antonym-synonym distinction. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 454–459, Berlin, Germany. Association for Computational Linguistics.
  16. 16.Xiaoman Pan, Kai Sun, Dian Yu, Jianshu Chen, Heng Ji, Claire Cardie, and Dong Yu. 2019. Improving question answering with external knowledge. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 27–37, Hong Kong, China. Association for Computational Linguistics.
  17. 17.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  18. 18.C. Riesbeck. 1981. Failure-driven reminding for incremental learning. In IJCAI.
  19. 19.Roger Schank. 1983. Dynamic Memory: A Theory of Reminding and Learning in Computers and People. Cambridge University Press.
  20. 20.Sida I. Wang, Percy Liang, and Christopher D. Manning. 2016. Learning language games through interaction. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2368–2378, Berlin, Germany. Association for Computational Linguistics.
  21. 21.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. ArXiv, abs/2109.01652.
  22. 22.Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. ArXiv, abs/2102.09690.

Citation

MLA
Madaan, A., et al. “Memory-assisted Prompt Editing to Improve GPT-3 After Deployment”. arXiv, 2022, http://arxiv.org/abs/2201.06009v7.
APA
Madaan, A., Tandon, N., Clark, P., & Yang, Y. (2022). Memory-assisted prompt editing to improve GPT-3 after deployment. arXiv. http://arxiv.org/abs/2201.06009v7
Chicago
Madaan, A., N. Tandon, P. Clark, and Y. Yang. 2022. “Memory-assisted Prompt Editing to Improve GPT-3 After Deployment”. arXiv. http://arxiv.org/abs/2201.06009v7.
Harvard
Madaan, A. et al. (2022) “Memory-assisted prompt editing to improve GPT-3 after deployment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.06009v7.
Vancouver
1. Madaan A, Tandon N, Clark P, Yang Y (2022) Memory-assisted prompt editing to improve GPT-3 after deployment. arXiv

BibTeX

@article{madaan2022memory,
  title = {Memory-assisted prompt editing to improve GPT-3 after deployment},
  author = {Madaan, Aman and Tandon, Niket and Clark, Peter and Yang, Yiming},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.06009v7},
  eprint = {2201.06009}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/