From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

Bernal Jimnez GutirrezYiheng ShuWeijian QiSizhe ZhouYu Su

article2025ICML255 citations

Proposes HippoRAG 2, a non-parametric continual learning framework that integrates knowledge graphs with Personalized PageRank and online language model reasoning to outperform standard retrieval-augmented generation across factual, sense-making, and associative memory tasks.

Listen

Large language models face severe limitations when absorbing new knowledge over time. Traditional continual fine-tuning causes models to catastrophically forget previously acquired information, while retrieval-augmented generation (RAG)—which supplies external data during inference—often fails at complex memory functions. Standard RAG relies on vector similarity searches that struggle with associative memory (connecting disparate facts across multi-hop questions) and sense-making (comprehending large, complex narratives). While newer structure-augmented systems attempt to address these gaps using knowledge graphs or summaries, they frequently inject noise that significantly degrades performance on basic factual retrieval.

The article introduces and evaluates HippoRAG 2, a non-parametric continual learning framework designed to approximate the dynamic, interconnected nature of human long-term memory. The framework aims to comprehensively improve retrieval and question-answering across factual recall, sense-making, and associative multi-hop reasoning without sacrificing performance on any single dimension.

The researchers developed HippoRAG 2 by integrating open knowledge extraction with the Personalized PageRank algorithm and three key architectural enhancements. First, the framework incorporates dense-sparse integration by adding passage nodes directly into the knowledge graph alongside conceptual phrase nodes. Second, it enhances query contextualization by matching queries to relational triples rather than isolated entity nodes. Third, it introduces a recognition memory filtering step powered by a language model to eliminate irrelevant triples before graph traversal. The authors evaluated the framework across seven diverse benchmark datasets—spanning simple factual question answering, multi-step associative reasoning, and full-length novel discourse comprehension—comparing it against standard dense embedding retrievers, large 7-billion-parameter embedding models, and multiple structure-augmented RAG frameworks.

The evaluation produced several critical findings. First, HippoRAG 2 achieved superior overall question-answering accuracy, leading the strongest baseline retriever by an average margin of 2.8 F1 points and outperforming standard graph-augmented baselines by over 10 points. Second, the system demonstrated exceptional associative reasoning, improving passage recall by 5.0% on MuSiQue and 13.9% on 2WikiMultihopQA compared to the leading dense retriever. Third, unlike competing graph-based and summarization methods, HippoRAG 2 maintained high factual precision and advanced sense-making on complex discourse without suffering performance drop-offs. Fourth, in continual learning simulations where corpus size expanded fourfold, HippoRAG 2 preserved its performance advantages consistently across both simple and multi-hop queries.

These results demonstrate that neurobiologically inspired retrieval architectures can effectively bridge the gap between simple text retrieval and human-like long-term memory. For organizations deploying artificial intelligence systems, HippoRAG 2 enables reliable non-parametric updates without the prohibitive costs, catastrophic forgetting risks, or fine-tuning overhead of parametric retraining. Furthermore, it operates effectively using both open-source models (such as Llama-3.3-70B) and proprietary models, providing operational flexibility and mitigating vendor lock-in.

Technical leaders and system architects should consider adopting triple-level contextual linking and graph-based reranking when building retrieval pipelines that demand multi-hop reasoning or long-document comprehension. In terms of trade-offs, HippoRAG 2 requires more graphical memory and slightly higher query latency (1.2 seconds per query) than pure dense retrieval (0.3 seconds), but it uses drastically fewer tokens and indexing hours than alternative community-summarization frameworks. Organizations should pilot this framework in high-stakes domains where complex reasoning outweighs modest computational overhead.

Confidence in these findings is bolstered by rigorous evaluations across varied dataset domains and retrieval backbones. However, decision-makers should note certain boundary conditions: error analyses revealed that 26% of retrieval failures stemmed from recognition filters discarding relevant triples or graph searches missing downstream nodes. Future development should focus on refining triple filtering precision and adapting graph-based non-parametric retrieval to enhance conversational episodic memory.

No sufficiently relevant recommendations were found.

Cover for From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

Abstract

Our ability to continuously acquire, organize, and leverage knowledge is a key feature of human intelligence that AI systems must approximate to unlock their full potential. Given the challenges in continual learning with large language models (LLMs), retrieval-augmented generation (RAG) has become the dominant way to introduce new information. However, its reliance on vector retrieval hinders its ability to mimic the dynamic and interconnected nature of human long-term memory. Recent RAG approaches augment vector embeddings with various structures like knowledge graphs to address some of these gaps, namely sense-making and associativity. However, their performance on more basic factual memory tasks drops considerably below standard RAG. We address this unintended deterioration and propose HippoRAG 2, a framework that outperforms standard RAG comprehensively on factual, sense-making, and associative memory tasks. HippoRAG 2 builds upon the Personalized PageRank algorithm used in HippoRAG and enhances it with deeper passage integration and more effective online use of an LLM. This combination pushes this RAG system closer to the effectiveness of human long-term memory, achieving a 7% improvement in associative memory tasks over the state-of-the-art embedding model while also exhibiting superior factual knowledge and sense-making memory capabilities. This work paves the way for non-parametric continual learning for LLMs. Code and data are available at https://github.com/OSU-NLP-Group/HippoRAG.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Continual Learning for LLMs
  • 2.2. Non-Parametric Continual Learning for LLMs
  • 3. HippoRAG 2
  • 3.1. Overview
  • 3.2. Dense-Sparse Integration
  • 3.3. Deeper Contextualization
  • 3.4. Recognition Memory
  • 3.5. Online Retrieval
  • 4. Experimental Setup
  • 4.1. Baselines
  • 4.2. Datasets
  • 4.3. Metrics
  • 4.4. Implementation Details
  • 5. Results
  • 6. Discussions
  • 6.1. Ablation Study
  • 6.2. Controlling Reset Probabilities
  • 6.3. Robustness to Corpus Expansion
  • 6.4. Dense Retriever Flexibility
  • 6.5. Qualitative Analysis
  • 7. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • Appendices
  • A. LLM Prompts
  • B. Pipeline Example
  • C. Detailed Experimental Results
  • D. Graph Statistics
  • E. Error Analysis
  • F. Cost and Efficiency
  • G. Implementation Details and Hyperparameters
  • G.1. HippoRAG 2
  • G.2. Comparison Methods

Knowls

  1. Knowl 1 — HippoRAG 2 combines graph-based association with passage-level context

    model/method

    HippoRAG 2 is a non-parametric retrieval-augmented memory system that constructs a knowledge graph (KG) offline and uses Personalized PageRank (PPR) for query-time passage retrieval. Its design integrates three kinds of information in retrieval: KG triples representing conceptual relations, passage nodes representing their source contexts, and the query’s embedding-based matches to triples and passages. An LLM filters candidate triples for query relevance before graph search; PPR then ranks passages for an LLM reader. The method is intended to improve associative and context-sensitive retrieval without sacrificing factual retrieval.

  2. Knowl 2 — Passage nodes integrate source context into the knowledge graph

    model/method

    HippoRAG 2 builds a schema-less KG by using an LLM for OpenIE on each corpus passage: each extracted triple contributes phrase nodes for its subject and object and a relation edge between them. An embedding model identifies phrase pairs whose vector similarity exceeds a threshold and connects them with synonym edges; the default threshold is 0.8. HippoRAG 2 then represents each original passage as a passage node and connects it to every phrase extracted from that passage with a context edge labeled contains. This addition links concise conceptual representations to the passages that provide their context, rather than using the KG to generate a separate summary corpus.

  3. Knowl 3 — Query-to-triple matching supplies contextual graph-search seeds

    model/method

    HippoRAG 2 links a query to KG facts by embedding the entire query and ranking KG triples by embedding similarity. It uses this query-to-triple approach by default, so the starting points for graph search can reflect relations among concepts as well as individual entity mentions. The paper compares it with two alternatives: NER-to-node extracts query entities and matches them to KG nodes, while query-to-node embeds the entire query and matches it directly to KG nodes. The authors motivate query-to-triple matching as a way to align query intent with richer contextual relations in the graph.

  4. Knowl 4 — Recognition memory filters candidate triples with an LLM

    model/method

    HippoRAG 2 models recognition memory as a relevance-filtering step after dense query-to-triple retrieval. The embedding model first returns the top five candidate triples. An LLM then selects up to four facts relevant to the query, may return an empty list, and is instructed to use only the candidates rather than inventing facts. The retained triples supply phrase nodes for graph search. In the paper’s implementation, Llama-3.3-70B-Instruct performs the filtering, with its prompt tuned using DSPy’s MIPROv2 optimizer.

  5. Knowl 5 — PPR retrieval uses both phrase and passage seed nodes

    algorithm

    HippoRAG 2 takes a query and an indexed KG as input and returns ranked corpus passages for downstream question answering. It uses normalized embedding scores for retrieval and initializes PPR as follows:

    1. Retrieve and LLM-filter the top five query-matched triples.
    2. If filtering leaves no usable phrase node, return the embedding-ranked passages directly, without graph search.
    3. Otherwise, select at most five phrase nodes from the retained triples. Score each phrase node by the average retrieval score of the filtered triples in which it appears.
    4. Use all passage nodes as seed nodes, with each passage’s embedding similarity as its score.
    5. Assign phrase nodes reset probabilities equal to their ranking scores. Multiply passage-node scores by 0.05 before assigning their reset probabilities, balancing the influence of phrase and passage seeds.
    6. Run PPR over the KG with damping factor 0.5, rank passage nodes by their resulting PageRank scores, and pass the top five passages to the QA reader.

    The 0.05 passage-node weight was selected based on validation results. The paper does not state a computational-complexity bound for this procedure.

  6. Knowl 6 — Evaluation covers factual, multi-hop, and long-discourse question answering

    experimental setup

    The evaluation measures factual memory with NaturalQuestions (NQ) and PopQA; associative, multi-hop retrieval with MuSiQue, 2WikiMultihopQA, HotpotQA, and LV-Eval; and long-discourse understanding with NarrativeQA. The sampled sets contain, respectively, 1,000 queries each for NQ, PopQA, MuSiQue, 2Wiki, and HotpotQA; all 124 LV-Eval queries; and 293 NarrativeQA queries. Their corpus sizes are 9,633, 8,676, 11,656, 6,119, 9,811, 22,849, and 4,111 passages, in that dataset order. Retrieval is evaluated with passage recall@5 and QA with token-based F1. In the main Llama-reader comparison, HippoRAG 2 and reproduced structure-augmented baselines use Llama-3.3-70B-Instruct for extraction and triple filtering and NV-Embed-v2 as retriever; the QA reader receives the top five retrieved passages. The paper also evaluates QA with GPT-4o-mini.

  7. Knowl 7 — HippoRAG 2 achieves the highest average QA F1 in the main comparison

    empirical result

    With Llama-3.3-70B-Instruct as the QA reader, HippoRAG 2 obtains an average F1 of 59.8 across the seven evaluated datasets, compared with 57.0 for NV-Embed-v2 and 53.1 for HippoRAG. In dataset order NQ, PopQA, MuSiQue, 2Wiki, HotpotQA, LV-Eval, and NarrativeQA, HippoRAG 2 scores 63.3, 56.2, 48.6, 71.0, 75.5, 12.9, and 25.9; NV-Embed-v2 scores 61.9, 55.7, 45.7, 61.5, 75.3, 9.8, and 25.7; and HippoRAG scores 55.3, 55.9, 35.1, 71.8, 63.5, 8.4, and 16.3. Thus, HippoRAG 2 has the highest reported average and improves notably on multi-hop and LV-Eval performance, although HippoRAG scores higher on 2Wiki and PopQA’s reported best score, 56.3, is marginally above HippoRAG 2’s 56.2. The paper marks HippoRAG 2’s NQ, MuSiQue, 2Wiki, and LV-Eval results as significantly better than the best NV-Embed-v2 baseline under a bootstrapped test at p < 0.05.

  8. Knowl 8 — HippoRAG 2 improves passage recall over dense retrieval

    empirical result

    On the five datasets with passage-retrieval annotations, HippoRAG 2 achieves an average passage recall@5 of 78.2, versus 73.4 for NV-Embed-v2, the strongest dense retriever in the comparison. In the order NQ, PopQA, MuSiQue, 2Wiki, and HotpotQA, HippoRAG 2 scores 78.0, 51.7, 74.7, 90.4, and 96.3; NV-Embed-v2 scores 75.4, 51.0, 69.7, 76.5, and 94.5. The largest gains over that dense baseline are 5.0 recall points on MuSiQue and 13.9 on 2Wiki. HippoRAG 2 also exceeds the reproduced HippoRAG on NQ, MuSiQue, and HotpotQA, ties it on 2Wiki at 90.4, and scores below it on PopQA (51.7 versus 53.8).

  9. Knowl 9 — Ablations show the contributions of contextual linking and graph components

    empirical result

    Passage recall@5 ablations on MuSiQue, 2Wiki, and HotpotQA show that the full HippoRAG 2 design performs best on average: 74.7, 90.4, and 96.3, respectively, for an average of 87.1. Replacing query-to-triple linking with NER-to-node yields 53.8, 91.2, and 78.8 (average 74.6); query-to-node yields 44.9, 65.5, and 68.3 (average 59.6). Removing passage nodes yields 63.7, 90.3, and 88.9 (average 81.0), while removing triple filtering yields 73.0, 90.7, and 95.4 (average 86.4). The results support the value of query-to-triple linking and passage nodes; filtering also improves the overall average, though its removal slightly increases recall on 2Wiki and HotpotQA.

    A separate validation-set test varied the passage-node reset-probability weight across 0.01, 0.05, 0.1, 0.3, and 0.5. MuSiQue recall@5 was 79.9, 80.5, 79.8, 78.4, and 77.9; NQ recall@5 was 75.6, 76.9, 76.9, 76.7, and 76.4. The authors selected 0.05 as the default weight.

  10. Knowl 10 — Performance gains persist as corpora grow and across dense retrievers

    empirical result

    In a corpus-expansion experiment, the authors divided NQ and MuSiQue into four segments, each containing gold documents and distractors for approximately 250 questions. They evaluated one segment while adding the other segments in stages, from 25% to 100% of the corpus. HippoRAG 2’s F1 advantage over NV-Embed-v2 remained broadly consistent on both factual NQ and multi-hop MuSiQue. Both systems retained strong NQ performance as documents were added, while their MuSiQue performance declined at a similar rate, showing that the associative task remained vulnerable to corpus growth.

    On a MuSiQue subset, HippoRAG 2 also outperformed direct dense retrieval with each tested encoder: GTE-Qwen2-7B-Instruct, 68.8 versus 63.6 recall@5; GritLM-7B, 71.6 versus 66.0; and NV-Embed-v2, 74.7 versus 69.7. In the GPT-4o-mini QA-reader comparison, HippoRAG 2’s average F1 was 58.1, compared with 55.7 for NV-Embed-v2.

Coverage note — The supplementary computational-cost comparison and error analysis were omitted: they characterize resource tradeoffs and failure cases but do not add another core method or benchmark result.

References

  1. 1.AI@Meta. Llama 3 model card. 2024. URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md.
  2. 2.Beyeler, M., Rounds, E. L., Carlson, K. D., Dutt, N., and Krichmar, J. L. Neural correlates of sparse coding and dimensionality reduction. PLoS Comput Biol, 15(6): e1006908, 2019. doi: 10.1371/journal.pcbi.1006908.
  3. 3.Chen, H., Pasunuru, R., Weston, J., and Celikyilmaz, A. Walking down the memory maze: Beyond context limit through interactive reading, 2023. URL https://arxiv.org/abs/2310.05029.
  4. 4.Cohen, R., Biran, E., Yoran, O., Globerson, A., and Geva, M. Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics, 12:283–298, 2024. doi: 10.1162/tacl a 00644. URL https://aclanthology.org/2024.tacl-1.16/.
  5. 5.Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., and Larson, J. From local to global: A graph rag approach to query-focused summarization, 2024. URL https://arxiv.org/abs/2404.16130.
  6. 6.Gu, J.-C., Xu, H.-X., Ma, J.-Y., Lu, P., Ling, Z.-H., Chang, K.-W., and Peng, N. Model editing harms general abilities of large language models: Regularization to the rescue. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 16801–16819, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.934. URL https://aclanthology.org/2024.emnlp-main.934/.
  7. 7.Guo, Z., Xia, L., Yu, Y., Ao, T., and Huang, C. LightRAG: Simple and fast retrieval-augmented generation, 2024. URL https://arxiv.org/abs/2410.05779.
  8. 8.Gutierrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., and Su, Y. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=hkujvAPVsg.
  9. 9.Haveliwala, T. H. Topic-sensitive pagerank. In Lassner, D., Roure, D. D., and Iyengar, A. (eds.), Proceedings of the Eleventh International World Wide Web Conference, WWW 2002, May 7-11, 2002, Honolulu, Hawaii, USA, pp. 517–526. ACM, 2002. doi: 10.1145/511446.511513. URL https://dl.acm.org/doi/10.1145/511446.511513.
  10. 10.Hoelscher-Obermaier, J., Persson, J., Kran, E., Konstas, I., and Barez, F. Detecting edit failures in large language models: An improved specificity benchmark. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Findings of the Association for Computational Linguistics: ACL 2023, pp. 11548–11559, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.733. URL https://aclanthology.org/2023.findings-acl.733/.
  11. 11.Huang, J., Cui, L., Wang, A., Yang, C., Liao, X., Song, L., Yao, J., and Su, J. Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1416–1428, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.77. URL https://aclanthology.org/2024.acl-long.77/.
  12. 12.Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E. Unsupervised dense information retrieval with contrastive learning. Trans. Mach. Learn. Res., 2022, 2022. URL https://openreview.net/forum?id=jKN1pXi7b0.
  13. 13.Jin, X., Zhang, D., Zhu, H., Xiao, W., Li, S.-W., Wei, X., Arnold, A., and Ren, X. Lifelong pretraining: Continually adapting language models to emerging corpora. In Carpuat, M., de Marneffe, M.-C., and Meza Ruiz, I. V. (eds.), Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4764–4780, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.351. URL https://aclanthology.org/2022.naacl-main.351/.
  14. 14.Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., Miller, H., Zaharia, M., and Potts, C. DSPy: Compiling declarative language model calls into self-improving pipelines. 2024.
  15. 15.Kim, Y., Yoon, J., Ye, S., Bae, S., Ho, N., Hwang, S. J., and Yun, S.-Y. Carpe diem: On the evaluation of world knowledge in lifelong language models. In Duh, K., Gomez, H., and Bethard, S. (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 5401–5415, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-long.302. URL https://aclanthology.org/2024.naacl-long.302/.
  16. 16.Klein, G., Moon, B., and Hoffman, R. R. Making sense of sensemaking 1: Alternative perspectives. IEEE intelligent systems, 21(4):70–73, 2006.
  17. 17.Koli, V., Yuan, J., and Dasgupta, A. Sensemaking of socially-mediated crisis information. In Blodgett, S. L., Cercas Curry, A., Dev, S., Madaio, M., Nenkova, A., Yang, D., and Xiao, Z. (eds.), Proceedings of the Third Workshop on Bridging Human–Computer Interaction and Natural Language Processing, pp. 74–81, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.hcinlp-1.7. URL https://aclanthology.org/2024.hcinlp-1.7/.
  18. 18.Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  19. 19.Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W. NV-embed: Improved techniques for training LLMs as generalist embedding models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=lgsyLSsDRe.
  20. 20.Li, J., Armandpour, M., Mirzadeh, S. I., Mehta, S., Shankar, V., Vemulapalli, R., Tuzel, O., Farajtabar, M., Pouransari, H., and Faghri, F. Tic-LM: A multi-year benchmark for continual pretraining of language models. In NeurIPS 2024 Workshop on Scalable Continual Learning for Lifelong Foundation Models, 2024. URL https://openreview.net/forum?id=PpSDVE5rAy.
  21. 21.Li, Z., Zhang, X., Zhang, Y., Long, D., Xie, P., and Zhang, M. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281, 2023.
  22. 22.Liska, A., Kocisky, T., Gribovskaya, E., Terzi, T., Sezener, E., Agrawal, D., De Masson D’Autume, C., Scholtes, T., Zaheer, M., Young, S., Gilsenan-Mcmahon, E., Austin, S., Blunsom, P., and Lazaridou, A. StreamingQA: A benchmark for adaptation to new knowledge over time in question answering models. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 13604–13622. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/liska22a.html.
  23. 23.Lu, X. H. BM25S: Orders of magnitude faster lexical search via eager sparse scoring, 2024. URL https://arxiv.org/abs/2407.03618.
  24. 24.Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9802–9822, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.546. URL https://aclanthology.org/2023.acl-long.546/.
  25. 25.Muennighoff, N., SU, H., Wang, L., Yang, N., Wei, F., Yu, T., Singh, A., and Kiela, D. Generative representational instruction tuning. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=BC4lIvfSzv.
  26. 26.Ni, J., Qu, C., Lu, J., Dai, Z., Hernandez Abrego, G., Ma, J., Zhao, V., Luan, Y., Hall, K., Chang, M.-W., and Yang, Y. Large dual encoders are generalizable retrievers. In Goldberg, Y., Kozareva, Z., and Zhang, Y. (eds.), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 9844–9855, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.669. URL https://aclanthology.org/2022.emnlp-main.669/.
  27. 27.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E. Z., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An imperative style, high-performance deep learning library. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 8024–8035, 2019.
  28. 28.Robertson, S. E. and Walker, S. Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval. In Croft, W. B. and van Rijsbergen, C. J. (eds.), Proceedings of the 17th Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval. Dublin, Ireland, 3-6 July 1994 (Special Issue of the SIGIR Forum), pp. 232–241. ACM/Springer, 1994. doi: 10.1007/978-1-4471-2099-5\24.
  29. 29.Roth, K., Udandarao, V., Dziadzio, S., Prabhu, A., Cherti, M., Vinyals, O., Henaff, O. J., Albanie, S., Bethge, M., and Akata, Z. A practitioner’s guide to continual multimodal pretraining. In NeurIPS 2024 Workshop on Scalable Continual Learning for Lifelong Foundation Models, 2024. URL https://openreview.net/forum?id=gkyosluSbR.
  30. 30.Sarthi, P., Abdullah, S., Tuli, A., Khanna, S., Goldie, A., and Manning, C. D. RAPTOR: recursive abstractive processing for tree-organized retrieval. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=GN921JHCRw.
  31. 31.Shi, H., Xu, Z., Wang, H., Qin, W., Wang, W., Wang, Y., Wang, Z., Ebrahimi, S., and Wang, H. Continual learning of large language models: A comprehensive survey. arXiv preprint arXiv:2404.16789, 2024.
  32. 32.Suzuki, W. A. Associative learning and the hippocampus. Psychological Science Agenda, February 2005.
  33. 33.Thakur, N., Reimers, N., Rucklé, A., Srivastava, A., and Gurevych, I. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=wCu6T5xFjeJ.
  34. 34.Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A. MuSiQue: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022. doi: 10.1162/tacl a 00475. URL https://aclanthology.org/2022.tacl-1.31/.
  35. 35.Uner, O. and Roediger III, H. L. Do recall and recognition lead to different retrieval experiences? The American Journal of Psychology, 135(1):33–43, 2022.
  36. 36.Wang, Y., Ren, R., Li, J., Zhao, X., Liu, J., and Wen, J. REAR: A relevance-aware retrieval-augmented framework for open-domain question answering. In Al-Onaizan, Y., Bansal, M., and Chen, Y. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pp. 5613–5626. Association for Computational Linguistics, 2024. URL https://aclanthology.org/2024.emnlp-main.321.
  37. 37.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771, 2019. URL http://arxiv.org/abs/1910.03771.
  38. 38.Xie, J., Zhang, K., Chen, J., Lou, R., and Su, Y. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=auKAUJZMO6.
  39. 39.Yao, Y., Wang, P., Tian, B., Cheng, S., Li, Z., Deng, S., Chen, H., and Zhang, N. Editing large language models: Problems, methods, and opportunities. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10222–10240, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.632. URL https://aclanthology.org/2023.emnlp-main.632/.
  40. 40.Yuan, T., Ning, X., Zhou, D., Yang, Z., Li, S., Zhuang, M., Tan, Z., Yao, Z., Lin, D., Li, B., Dai, G., Yan, S., and Wang, Y. LV-Eval: A balanced long-context benchmark with 5 length levels up to 256k, 2024. URL https://arxiv.org/abs/2402.05136.
  41. 41.Zhang, H., Gui, L., Zhai, Y., Wang, H., Lei, Y., and Xu, R. Copr: Continual learning human preference through optimal policy regularization, 2024. URL https://arxiv.org/abs/2310.15694.
  42. 42.Zhang, Z., Fang, M., Chen, L., and Namazi-Rad, M.-R. CITB: A benchmark for continual instruction tuning. In Bouamor, H., Pino, J., and Bali, K. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 9443–9455, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.633. URL https://aclanthology.org/2023.findings-emnlp.633/.
  43. 43.Zhong, Z., Wu, Z., Manning, C., Potts, C., and Chen, D. MQuAKE: Assessing knowledge editing in language models via multi-hop questions. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 15686–15702, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.971. URL https://aclanthology.org/2023.emnlp-main.971/.

Citation

MLA
Gutiérrez, B. J., et al. “From RAG to Memory: Non-Parametric Continual Learning for Large Language Models”. arXiv, 2025, http://arxiv.org/abs/2502.14802v2.
APA
Gutiérrez, B. J., Shu, Y., Qi, W., Zhou, S., & Su, Y. (2025). From RAG to Memory: Non-Parametric Continual Learning for Large Language Models. arXiv. http://arxiv.org/abs/2502.14802v2
Chicago
Gutiérrez, B. J., Y. Shu, W. Qi, S. Zhou, and Y. Su. 2025. “From RAG to Memory: Non-Parametric Continual Learning for Large Language Models”. arXiv. http://arxiv.org/abs/2502.14802v2.
Harvard
Gutiérrez, B.J. et al. (2025) “From RAG to Memory: Non-Parametric Continual Learning for Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.14802v2.
Vancouver
1. Gutiérrez BJ, Shu Y, Qi W, Zhou S, Su Y (2025) From RAG to Memory: Non-Parametric Continual Learning for Large Language Models. arXiv

BibTeX

@article{gutierrez2025from,
  title = {From RAG to Memory: Non-Parametric Continual Learning for Large Language Models},
  author = {Gutiérrez, Bernal Jiménez and Shu, Yiheng and Qi, Weijian and Zhou, Sizhe and Su, Yu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.14802v2},
  eprint = {2502.14802}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/