MEMORYLLM: Towards Self-Updatable Large Language Models

Yu WangYifan GaoXiusi ChenHaoming JiangShiyang LiJingfeng YangQingyu YinZheng LiXian LiBing Yin

article2024ICML50 citations

Introduces MemoryLLM, an architecture that embeds a fixed-size, self-updatable latent memory pool into transformer layers to continuously absorb new textual knowledge and retain past information across nearly a million update cycles without degrading model capabilities.

Listen

Modern large language models are generally static once trained and deployed, making it challenging, computationally expensive, and inefficient to incorporate new facts or update outdated information. Conventional workarounds—such as external retrieval systems, extended prompt contexts, and targeted weight-editing methods—often suffer from uncontrolled storage growth, severe computational bottlenecks, or unintended distortion of existing knowledge.

The article introduces and evaluates MemoryLLM, an architecture that embeds a fixed-capacity, self-updatable memory pool directly within the latent space of a language model. The main objective is to demonstrate that an artificial intelligence model can continuously integrate new knowledge without retraining, while retaining older information through a controlled, exponential forgetting process and maintaining long-term operational integrity.

To evaluate this approach, the authors augmented a standard 7-billion-parameter model with an internal memory pool comprising roughly 1 billion self-updating parameters distributed across all model layers. During updates, the model integrates new text by modifying a small portion of the memory and randomly replacing older memory tokens, eliminating the need for expensive back-propagation during deployment. The framework was evaluated across standardized model editing benchmarks, long-context question-answering tasks, dedicated knowledge retention experiments, and extended stress tests spanning hundreds of thousands of updates.

The key findings demonstrate clear advantages over conventional methods. On model editing benchmarks, MemoryLLM achieved superior overall performance scores of 79.2 and 75.3 on standard test sets, significantly outperforming existing editing techniques while preserving broad factual accuracy. On long-context benchmarks, the model outperformed baseline systems on four out of six evaluation datasets, maintaining high accuracy as input length scaled while operating within standard hardware memory limits. Furthermore, in long-term retention experiments, the model retained retrievable knowledge across dozens of consecutive updates, closely tracking theoretical exponential decay models. Finally, in extreme operational stress testing involving 650,000 continuous update cycles, the architecture exhibited zero performance degradation, verifying that repeated updates do not destabilize the underlying model.

These results indicate that internal, fixed-size memory pools offer a practical path toward continuously updatable artificial intelligence. For organizations deploying language models, this approach significantly reduces the operational overhead, infrastructure costs, and latency associated with continuous retraining or managing ever-expanding retrieval databases. It strikes an effective balance between rapid knowledge absorption and controlled, predictable phase-out of stale information.

Decision-makers considering dynamic knowledge architectures should monitor the development of self-updating memory frameworks as a viable alternative or complement to standard retrieval pipelines. Before broad deployment, organizations should conduct domain-specific pilot testing, scale the memory architecture to larger foundation models, and investigate multimodal extensions. Users should note that performance on specialized technical texts showed sensitivity to pre-training domain coverage, requiring careful alignment with target enterprise domains.

  • Paper: Memorizing Transformers, Yuhuai Wu et al. (2022). It introduces the foundational architectural concept of augmenting transformer layers with an explicit, non-differentiable memory cache of past token states to extend model capacity without back-propagation.
  • Book: AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models, Junfeng Fang et al. (2025). It details the core challenge and baseline methodologies of sequential model editing and preserving factual knowledge without catastrophic interference, motivating the self-updatable latent memory pool in MemoryLLM.
  • Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). It establishes essential benchmark protocols and evaluation methodologies for sequential knowledge editing and factual updates in language models.
  • Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). It provides a key system design paradigm for treating LLM memory hierarchically to overcome context limits, establishing the functional context for internal latent memory mechanisms.
  • Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). It explores parameter-free core continual learning via external prompt memory pools, offering foundational principles for updating model behavior without full parameter retraining.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It outlines the foundational theoretical framework for reading and writing to explicit memory stores integrated end-to-end with neural networks.
  • Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It provides a comprehensive theoretical taxonomy of continual learning and stability-plasticity trade-offs that underpin the controlled forgetting dynamics of MemoryLLM.
  • Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). It extends the paradigm of dynamic model memory by training a dedicated, modular memory model to parametrically supply up-to-date knowledge to frozen executive LLMs.
  • Paper: $δ$-mem: Efficient Online Memory for Large Language Models, Jingdi Lei et al. (2026). It advances lightweight online memory updates by integrating an evolving, low-parameter state into attention mechanisms during multi-turn interactions.
  • Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). It develops neural test-time memory architectures with gradient-based surprise metrics and adaptive forgetting gates to dynamically absorb long-range context.
  • Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). It expands continual test-time learning directly into standard transformer sub-layers to achieve scalable context compression and memorization during inference.
  • Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). It investigates multi-mechanism continual learning compositions for long-horizon memorization and retention under extensive sequential updates.
  • Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It provides a dedicated evaluation benchmark to assess how well dynamic memory and long-context systems retain temporal consistency and factual information over extended deployment periods.
Cover for MEMORYLLM: Towards Self-Updatable Large Language Models

Abstract

Existing Large Language Models (LLMs) usually remain static after deployment, which might make it hard to inject new knowledge into the model. We aim to build models containing a considerable portion of self-updatable parameters, enabling the model to integrate new knowledge effectively and efficiently. To this end, we introduce MEMORYLLM, a model that comprises a transformer and a fixed-size memory pool within the latent space of the transformer. MEMORYLLM can self-update with text knowledge and memorize the knowledge injected earlier. Our evaluations demonstrate the ability of MEMORYLLM to effectively incorporate new knowledge, as evidenced by its performance on model editing benchmarks. Meanwhile, the model exhibits long-term information retention capacity, which is validated through our custom-designed evaluations and long-context benchmarks. MEMORYLLM also shows operational integrity without any sign of performance degradation even after nearly a million memory updates. Our code and model are open-sourced at https://github.com/wangyu-ustc/MemoryLLM.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Problem Statement
  • 2.2. Sketch of MEMORYLLM
  • 3. MEMORYLLM
  • 3.1. Structure Design
  • 3.1.1. MEMORY POOL
  • 3.1.2. SELF-UPDATE PROCESS
  • 3.1.3. ANALYSIS OF FORGETTING
  • 3.2. Training Strategy
  • 3.2.1. NEW KNOWLEDGE INCORPORATION
  • 3.2.2. ENHANCING CONTINUOUS CONTEXTS UNDERSTANDING
  • 3.2.3. MITIGATING FORGETTING PROBLEMS
  • 3.3. Model Instantiation
  • 3.4. Discussions
  • 4. Experiments
  • 4.1. Evaluation Protocols
  • 4.2. Implementation Details
  • 4.3. Model Editing
  • 4.3.2. OVERALL PERFORMANCE COMPARISON
  • 4.4. Long Context Evaluation
  • 4.4.1. EXPERIMENTAL SETUP
  • 4.4.2. OVERALL PERFORMANCE COMPARISON
  • 4.4.3. COMPARISON WITH RAG METHODS
  • 4.5. Knowledge Retention Experiments
  • 4.5.1. EXPERIMENTAL SETUP
  • 4.6. Model Integrity Analysis
  • 4.7. Ablation Study
  • 4.7.1. ABLATION STUDY OF DIFFERENT K AND N
  • 4.7.2. ABLATION STUDY OF THE MODEL STRUCTURES
  • 5. Related Work
  • 5.1. Memory based methods
  • 5.2. Downsteam Tasks
  • 6. Conclusion and Future Work
  • Impact Statement
  • References
  • A. Details in Methodology
  • A.1. Self-Update Process
  • A.2. Training Strategy for New Knowledge Incorporation
  • B. Implementation Details
  • B.1. Details for Mitigating Forgetting Problems
  • C. Additional Experiments
  • C.1. Baselines for Model Editing

Knowls

  1. Knowl 1 — MemoryLLM Model Architecture and Latent Memory Pool

    model/method

    MemoryLLM partitions model parameters into static parameters ϕ\phi and dynamically self-updatable parameters θ\theta. The static component ϕ\phi is instantiated as a standard multi-layer transformer with LL layers, ϕ={ϕl}l=1L\phi = \{\phi_l\}_{l=1}^L. The dynamic component θ\theta is structured as a fixed-size latent memory pool embedded across all layers of the transformer, represented as θ={θl}l=1L\theta = \{\theta^l\}_{l=1}^L.

    Each layer's memory θl∈RN×d\theta^l \in \mathbb{R}^{N \times d} consists of NN memory tokens of embedding dimension dd, where dd matches the transformer hidden dimension. In the primary implementation based on Llama2-7B (L=32L = 32, d=4096d = 4096), each layer is allocated N=7680N = 7680 memory tokens, yielding a total memory parameter count of 32×7680×4096≈1.066B32 \times 7680 \times 4096 \approx 1.066\text{B} parameters.

    During the generation phase, an input text sequence xx comprising nxn_x tokens produces hidden states hlh_l at layer ll. The attention mechanism allows every token in hlh_l to attend to all NN memory tokens in θl\theta^l as well as all preceding sequence tokens, creating an attention map of dimension nx×(nx+N)n_x \times (n_x + N). This formulation yields a computational complexity that scales linearly with the memory pool token count NN for each generated token.

  2. Knowl 2 — Self-Update Mechanism via Partial Memory Extraction and Random Token Dropping

    model/method

    MemoryLLM integrates new text context xcx_c of length nxcn_{x_c} into the memory pool θ\theta through a feedforward self-update function θ′=U(θ,xc)\theta' = U(\theta, x_c) without requiring gradient backpropagation during deployment.

    For each transformer layer l∈{1,…,L}l \in \{1, \dots, L\}:

    1. The most recent KK memory tokens (where K≪NK \ll N) are extracted from the current layer memory θl\theta^l, denoted as eθl∈RK×de_\theta^l \in \mathbb{R}^{K \times d}.

    2. The extracted memory tokens eθle_\theta^l are concatenated with the sequence hidden states hl∈Rnxc×dh_l \in \mathbb{R}^{n_{x_c} \times d} (where h1=Embed(xc)h_1 = \text{Embed}(x_c)). The concatenated input is fed into transformer layer ϕl\phi_l using an attention map of dimension max⁡(nxc,K)×(nxc+K)\max(n_{x_c}, K) \times (n_{x_c} + K), allowing the hidden states hlh_l to attend to the prior memory tokens eθle_\theta^l.

    3. The last KK hidden vectors of the layer output are extracted to serve as the updated memory tokens eθl′∈RK×de_\theta^{l\prime} \in \mathbb{R}^{K \times d}, while the contextual representations form hl+1h_{l+1} for the next layer.

    4. A subset of KK memory tokens is selected uniformly at random from the existing pool θl\theta^l and discarded, leaving N−KN - K retained memory tokens denoted θl(d)\theta^l(d).

    5. The retained tokens are shifted and concatenated with the new memory tokens to form the new layer memory θl′=[θl(d);eθl′]∈RN×d\theta^{l\prime} = [\theta^l(d); e_\theta^{l\prime}] \in \mathbb{R}^{N \times d}, maintaining a strictly constant memory size NN.

  3. Knowl 3 — Theoretical Exponential Forgetting Rate in Fixed-Capacity Memory

    theoretical result

    Because MemoryLLM removes KK tokens chosen uniformly at random from the total pool of NN memory tokens at every self-update step, the probability of any given memory token being retained across a single update is 1−KN1 - \frac{K}{N}.

    Under this dropping mechanism, the proportion of information remaining from knowledge injected t−1t-1 update steps prior follows an exponential decay:

    R(t)=(1−KN)t−1R(t) = \left(1 - \frac{K}{N}\right)^{t-1}

    After NK\frac{N}{K} update steps, the theoretical retention limit as the memory pool scales satisfies:

    lim⁡NK→∞(1−KN)NK=1e≈0.3679\lim_{\frac{N}{K} \to \infty} \left(1 - \frac{K}{N}\right)^{\frac{N}{K}} = \frac{1}{e} \approx 0.3679

    Given an initial prediction accuracy aua_u achieved immediately after injection (at step 1) and a baseline borderline accuracy aba_b without context injection, the theoretical upper-bound prediction accuracy ata_t after t−1t-1 subsequent updates is:

    at=(au−ab)(1−KN)t−1+aba_t = (a_u - a_b) \left(1 - \frac{K}{N}\right)^{t-1} + a_b

  4. Knowl 4 — Decomposed Gradient Training and Memory Regularization

    model/method

    Pretraining MemoryLLM on next-token prediction uses three combined training routines to enable knowledge assimilation while managing GPU memory overhead:

    1. New Knowledge Incorporation with Decomposed Gradients: A document dd is divided into consecutive segments (x1,x2)(x_1, x_2). Training iterates randomly with 50% probability between two computational pathways:

      • Gradient pathway: Context x1x_1 is self-updated through the model with autograd enabled, but only the resulting new memory tokens eθl′∈RK×de_\theta^{l\prime} \in \mathbb{R}^{K \times d} (rather than the full pool θl′\theta^{l\prime}) are concatenated as context to compute cross-entropy loss on x2x_2. This limits peak memory during backpropagation while optimizing knowledge compression from x1x_1 into eθl′e_\theta^{l\prime}.
      • Non-gradient pathway: Memory θ\theta is updated on x1x_1 with autograd disabled, and the entire updated memory pool θ′\theta' is used to predict x2x_2.
    2. Continuous Context Understanding: Documents exceeding 2048 tokens are partitioned into nn segments (x1,…,xn)(x_1, \dots, x_n). Segments x1,…,xn−1x_1, \dots, x_{n-1} are sequentially injected into θ\theta with autograd disabled according to θk=U(θk−1,xk)\theta_k = U(\theta_{k-1}, x_k). Cross-entropy loss is then computed on the final segment xnx_n conditioned on θn−1\theta_{n-1}.

    3. Distribution Integrity Regularization: At the end of each training iteration, the memory pool θ\theta is permanently updated with the processed text inputs (x1x_1 or {x1,…,xn−1}\{x_1, \dots, x_{n-1}\}) post-backpropagation. This aligns the underlying vector distribution of θl\theta^l with newly generated eθl′e_\theta^{l\prime}, preventing drift over infinite updates.

  5. Knowl 5 — Interleaved Training Algorithm for Long-Term Knowledge Retention

    algorithm

    To train the model to retrieve facts injected multiple steps in the past and mitigate catastrophic forgetting, training interleaves segments across different documents using a randomized cache strategy:

    Input: Training dataset DD
    Initialize indicator r0=1r_0 = 1, counter l=0l = 0, cached target segment xcache=Nonex_{\text{cache}} = \text{None}
    for document d∈Dd \in D do
        n=number of segments in dn = \text{number of segments in } d
        {x1,x2,…,xn}=d\{x_1, x_2, \dots, x_n\} = d
        if r0=1r_0 = 1 or l=0l = 0 then
            r=0r = 0
        else
            r∼UniformRandom({0,1})r \sim \text{UniformRandom}(\{0, 1\})
        end if
        if r=0r = 0 and r0=0r_0 = 0 then
            Inject {x1,…,xn−1}\{x_1, \dots, x_{n-1}\} into memory pool θ\theta without gradients
            Calculate cross-entropy loss on segment xnx_n and update model parameters ϕ\phi
            l=l+nl = l + n
        else if r=0r = 0 and r0=1r_0 = 1 then
            Inject {x1,…,xn−1}\{x_1, \dots, x_{n-1}\} into memory pool θ\theta without gradients
            Calculate cross-entropy loss on segment xnx_n and update model parameters ϕ\phi
            xcache=xnx_{\text{cache}} = x_n
            l=l+nl = l + n
        else if r=1r = 1 then
            Calculate cross-entropy loss on cached segment xcachex_{\text{cache}} and update model parameters ϕ\phi
            l=0l = 0
        end if
        r0=rr_0 = r
    end for

    This procedure forces the model to recall and predict xcachex_{\text{cache}} after an arbitrary number of intervening contexts from other documents have updated the memory pool.

  6. Knowl 6 — Model Editing Performance on zsRE and CounterFactual Benchmarks

    data/table

    MemoryLLM is evaluated on fact editing without gradient updates during inference against standard model editing techniques. Evaluations use the first 10,000 records of Zero-Shot Relation Extraction (zsRE) and 2,000 records of CounterFactual. Evaluated metrics are Efficacy (post-edit accuracy on target fact), Generalization (post-edit accuracy on paraphrased fact), Specificity (post-edit accuracy on unrelated facts), and Score (harmonic mean of the three metrics).

    ZsRE Dataset CounterFactual Dataset
    Editor Score Efficacy Generalization Specificity Score Efficacy Generalization Specificity
    Llama2-7B 55.6 55.9 54.7 56.3 20.7 13.7 16.6 83.4
    MemoryLLM-7B 51.2 50.0 49.1 54.8 22.6 15.7 17.6 82.1
    FT 50.3 78.6 80.6 29.0 10.0 99.7 96.9 3.6
    FT-L 69.8 81.4 76.8 56.6 33.8 47.2 18.0 83.3
    ROME 69.3 88.7 70.2 56.3 69.2 82.6 75.2 55.8
    IKE - - - - 70.7 99.8 96.2 45.4
    MemoryLLM-7B (w/ EF) 79.2 99.8 96.7 57.1 75.3 98.5 82.2 57.0

    MemoryLLM-7B with editing facts injected into memory (w/ EF) achieves the highest harmonic Score across both datasets (79.2 on zsRE and 75.3 on CounterFactual). While standard fine-tuning (FT) attains high Efficacy (99.7%) but suffers catastrophic collapse on Specificity (3.6%), and constrained fine-tuning (FT-L) preserves Specificity (83.3%) at the expense of Efficacy (47.2%), MemoryLLM maintains high scores across Efficacy (98.5%), Generalization (82.2%), and Specificity (57.0%) simultaneously.

  7. Knowl 7 — Long-Context Question Answering Performance and Hardware Efficiency on LongBench

    empirical result

    MemoryLLM-7B was tested across context lengths ranging from 2k to 16k tokens on six LongBench datasets: NarrativeQA, Qasper, MultiFieldQA-en, HotpotQA, 2WikiMultiHopQA, and MuSiQue. The model was compared against un-instruction-tuned baselines: OpenLlama-3B-v2, LongLlama-3B-v1.1, Llama2-7B, Llama2-LongLora-7B-16k, and Llama2-LongLora-7B-100k.

    1. Performance: MemoryLLM-7B outperforms all baseline models on 4 out of 6 datasets (NarrativeQA, MultiFieldQA-en, HotpotQA, and 2WikiMultiHopQA) as context length increases to 16k. Performance on Qasper was suboptimal, attributable to training MemoryLLM primarily on C4 rather than scientific arXiv text.

    2. Computational and Memory Footprint: Standard 7B baseline models (such as Llama2-LongLora-7B) encountered Out-of-Memory (OOM) errors at context lengths of 16,384 even when using eight 80GB A100 GPUs. In contrast, because MemoryLLM processes long context via sequential self-updates of size KK into a fixed-size latent memory, it requires only a single 48GB GPU or two 40GB GPUs to run inference regardless of the input text length.

  8. Knowl 8 — Operational Integrity and Stability of MemoryLLM Over 650,000 Continuous Updates

    empirical result

    To evaluate whether continuous latent updates cause numerical instability, memory drift, or model breakdown, MemoryLLM was subjected to 650,000 sequential update steps over 3 continuous days by repeatedly shuffling and injecting QA contexts from SQuAD (2,250 examples) and NaturalQA (1,004 examples).

    The accuracy of the model on the most recently injected context was recorded at every step. The exponentially smoothed accuracy (with a smoothing factor of 99.99%) remained completely flat without any sign of degradation across all 650,000 updates, staying centered near 0.38 on SQuAD and 0.46 on NaturalQA. This demonstrates that updating memory tokens post-backpropagation during training regularizes latent activations and guarantees long-term operational integrity under infinite self-updates.

  9. Knowl 9 — Ablations on Memory Capacity $N$, Update Compression Size $K$, and Layer Distribution

    empirical result

    Ablation experiments on SQuAD and NaturalQA evaluate the impact of memory parameters NN, KK, and architectural placement on retention:

    1. Memory Capacity NN and Update Size KK: Testing configurations with N∈{10×256,20×256,30×256}N \in \{10 \times 256, 20 \times 256, 30 \times 256\} and K∈{256,512}K \in \{256, 512\} demonstrates that the retention ratio improves as the ratio N/KN/K increases. For a fixed K=256K = 256, higher NN (30×25630 \times 256) produces the slowest forgetting curve across 15 update steps. For a fixed total memory size of 5120 tokens, K=256K = 256 (20×25620 \times 256) outperforms K=512K = 512 (10×51210 \times 512). Reducing KK to 128 sharply impairs step-1 accuracy (0.34 vs 0.46 on NaturalQA; 0.25 vs 0.39 on SQuAD), indicating an information compression limit.

    2. Layer Allocation: Augmenting only a single transformer layer with memory tokens yields almost zero contextual prediction improvement over a baseline without context. Augmenting only the top half of transformer layers (layers 17–32) achieves lower step-1 accuracy (0.39 on NaturalQA, 0.22 on SQuAD) compared to placing memory tokens across all 32 layers (0.46 on NaturalQA, 0.39 on SQuAD).

    3. Token Dropping vs. Memory Aggregation: Replacing random token dropping with decaying exponential aggregation degrades performance because vector blending corrupts both existing and newly injected hidden states, preventing clean extraction of recently learned facts.

Coverage note — No substantial contributed material was omitted. The orthogonal experiment evaluating BM25 retriever combination with MemoryLLM (Section 4.4.3, Table 2) was integrated into the contextual discussions.

References

  1. 1.Adel, A. A. Global memory transformer for processing long documents. CoRR, abs/2212.01650, 2022.
  2. 2.Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508, 2023.
  3. 3.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. CoRR, abs/2004.05150, 2020. URL https://arxiv.org/abs/2004.05150.
  4. 4.Bulatov, A., Kuratov, Y., and Burtsev, M. S. Recurrent memory transformer. In NeurIPS, 2022.
  5. 5.Burtsev, M. S. and Sapunov, G. V. Memory transformer. CoRR, abs/2006.11527, 2020. URL https://arxiv.org/abs/2006.11527.
  6. 6.Cao, N. D., Aziz, W., and Titov, I. Editing factual knowledge in language models. In EMNLP (1), pp. 6491–6506. Association for Computational Linguistics, 2021.
  7. 7.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023a.
  8. 8.Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J. Longlora: Efficient fine-tuning of long-context large language models. arXiv preprint arXiv:2309.12307, 2023b.
  9. 9.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. CoRR, abs/1904.10509, 2019. URL http://arxiv.org/abs/1904.10509.
  10. 10.Computer, T. Redpajama: an open dataset for training large language models, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  11. 11.Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. Meshed-memory transformer for image captioning. In CVPR, pp. 10575–10584. Computer Vision Foundation / IEEE, 2020.
  12. 12.Fang, J., Tang, L., Bi, H., Qin, Y., Sun, S., Li, Z., Li, H., Li, Y., Cong, X., Yan, Y., et al. Unimem: Towards a unified view of long-context large language models. arXiv preprint arXiv:2402.03009, 2024.
  13. 13.Geng, X. and Liu, H. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama.
  14. 14.Jiayu, D., Shuming, M., Li, D., Xingxing, Z., Shaohan, H., Wenhui, W., and Wei†, F. Longnet: Scaling transformers to 1,000,000,000 tokens. arXiv preprint arXiv:2307.02486, 2023.
  15. 15.Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. Generalization through memorization: Nearest neighbor language models. arXiv preprint arXiv:1911.00172, 2019.
  16. 16.Levy, O., Seo, M., Choi, E., and Zettlemoyer, L. Zero-shot relation extraction via reading comprehension. In Levy, R. and Specia, L. (eds.), Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Vancouver, Canada, August 3-4, 2017, pp. 333–342. Association for Computational Linguistics, 2017. doi: 10.18653/V1/K17-1034. URL https://doi.org/10.18653/v1/K17-1034.
  17. 17.Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372, 2022.
  18. 18.Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D. Fast model editing at scale. In ICLR. OpenReview.net, 2022.
  19. 19.Modarressi, A., Imani, A., Fayyaz, M., and Schütze, H. Ret-llm: Towards a general read-write memory for large language models. arXiv preprint arXiv:2305.14322, 2023.
  20. 20.Moro, G., Ragazzi, L., Valgimigli, L., Frisoni, G., Sartori, C., and Marfia, G. Efficient memory-enhanced transformer for long-document summarization in low-resource regimes. Sensors, 23(7):3542, 2023.
  21. 21.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021.
  22. 22.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020.
  23. 23.Sukhbaatar, S., Weston, J., Fergus, R., et al. End-to-end memory networks. Advances in neural information processing systems, 28, 2015.
  24. 24.Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F. A length-extrapolatable transformer. In ACL (1), pp. 14590–14604. Association for Computational Linguistics, 2023.
  25. 25.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023.
  26. 26.Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłos, P. Focused transformer: Contrastive training for context scaling. arXiv preprint arXiv:2307.03170, 2023.
  27. 27.Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  28. 28.Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F. Augmenting language models with long-term memory. arXiv preprint arXiv:2306.07174, 2023.
  29. 29.Weston, J., Chopra, S., and Bordes, A. Memory networks. arXiv preprint arXiv:1410.3916, 2014.
  30. 30.Wu, Q., Lan, Z., Qian, K., Gu, J., Geramifard, A., and Yu, Z. Memformer: A memory-augmented transformer for sequence modeling. In AACL/IJCNLP (Findings), pp. 308–318. Association for Computational Linguistics, 2022.
  31. 31.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al. Effective long-context scaling of foundation models. arXiv preprint arXiv:2309.16039, 2023.
  32. 32.Yao, Y., Wang, P., Tian, B., Cheng, S., Li, Z., Deng, S., Chen, H., and Zhang, N. Editing large language models: Problems, methods, and opportunities. CoRR, abs/2305.13172, 2023.
  33. 33.Zheng, C., Li, L., Dong, Q., Fan, Y., Wu, Z., Xu, J., and Chang, B. Can we edit factual knowledge by in-context learning? arXiv preprint arXiv:2305.12740, 2023.
  34. 34.Zhong, W., Guo, L., Gao, Q., and Wang, Y. Memorybank: Enhancing large language models with long-term memory. arXiv preprint arXiv:2305.10250, 2023.
  35. 35.Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F. X., and Kumar, S. Modifying memories in transformer models. CoRR, abs/2012.00363, 2020.

Citation

MLA
Wang, Y., et al. “MEMORYLLM: Towards Self-Updatable Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.04624v2.
APA
Wang, Y., Gao, Y., Chen, X., Jiang, H., Li, S., Yang, J., Yin, Q., Li, Z., Li, X., Yin, B., Shang, J., & McAuley, J. (2024). MEMORYLLM: Towards Self-Updatable Large Language Models. arXiv. http://arxiv.org/abs/2402.04624v2
Chicago
Wang, Y., Y. Gao, X. Chen, et al. 2024. “MEMORYLLM: Towards Self-Updatable Large Language Models”. arXiv. http://arxiv.org/abs/2402.04624v2.
Harvard
Wang, Y. et al. (2024) “MEMORYLLM: Towards Self-Updatable Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.04624v2.
Vancouver
1. Wang Y, Gao Y, Chen X, et al (2024) MEMORYLLM: Towards Self-Updatable Large Language Models. arXiv

BibTeX

@article{wang2024memoryllm,
  title = {MEMORYLLM: Towards Self-Updatable Large Language Models},
  author = {Wang, Yu and Gao, Yifan and Chen, Xiusi and Jiang, Haoming and Li, Shiyang and Yang, Jingfeng and Yin, Qingyu and Li, Zheng and Li, Xian and Yin, Bing and Shang, Jingbo and McAuley, Julian},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.04624v2},
  eprint = {2402.04624}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/