MEMORYLLM: Towards Self-Updatable Large Language Models
Yu WangYifan GaoXiusi ChenHaoming JiangShiyang LiJingfeng YangQingyu YinZheng LiXian LiBing Yin
Introduces MemoryLLM, an architecture that embeds a fixed-size, self-updatable latent memory pool into transformer layers to continuously absorb new textual knowledge and retain past information across nearly a million update cycles without degrading model capabilities.
Modern large language models are generally static once trained and deployed, making it challenging, computationally expensive, and inefficient to incorporate new facts or update outdated information. Conventional workarounds—such as external retrieval systems, extended prompt contexts, and targeted weight-editing methods—often suffer from uncontrolled storage growth, severe computational bottlenecks, or unintended distortion of existing knowledge.
The article introduces and evaluates MemoryLLM, an architecture that embeds a fixed-capacity, self-updatable memory pool directly within the latent space of a language model. The main objective is to demonstrate that an artificial intelligence model can continuously integrate new knowledge without retraining, while retaining older information through a controlled, exponential forgetting process and maintaining long-term operational integrity.
To evaluate this approach, the authors augmented a standard 7-billion-parameter model with an internal memory pool comprising roughly 1 billion self-updating parameters distributed across all model layers. During updates, the model integrates new text by modifying a small portion of the memory and randomly replacing older memory tokens, eliminating the need for expensive back-propagation during deployment. The framework was evaluated across standardized model editing benchmarks, long-context question-answering tasks, dedicated knowledge retention experiments, and extended stress tests spanning hundreds of thousands of updates.
The key findings demonstrate clear advantages over conventional methods. On model editing benchmarks, MemoryLLM achieved superior overall performance scores of 79.2 and 75.3 on standard test sets, significantly outperforming existing editing techniques while preserving broad factual accuracy. On long-context benchmarks, the model outperformed baseline systems on four out of six evaluation datasets, maintaining high accuracy as input length scaled while operating within standard hardware memory limits. Furthermore, in long-term retention experiments, the model retained retrievable knowledge across dozens of consecutive updates, closely tracking theoretical exponential decay models. Finally, in extreme operational stress testing involving 650,000 continuous update cycles, the architecture exhibited zero performance degradation, verifying that repeated updates do not destabilize the underlying model.
These results indicate that internal, fixed-size memory pools offer a practical path toward continuously updatable artificial intelligence. For organizations deploying language models, this approach significantly reduces the operational overhead, infrastructure costs, and latency associated with continuous retraining or managing ever-expanding retrieval databases. It strikes an effective balance between rapid knowledge absorption and controlled, predictable phase-out of stale information.
Decision-makers considering dynamic knowledge architectures should monitor the development of self-updating memory frameworks as a viable alternative or complement to standard retrieval pipelines. Before broad deployment, organizations should conduct domain-specific pilot testing, scale the memory architecture to larger foundation models, and investigate multimodal extensions. Users should note that performance on specialized technical texts showed sensitivity to pre-training domain coverage, requiring careful alignment with target enterprise domains.
- Paper: Memorizing Transformers, Yuhuai Wu et al. (2022). It introduces the foundational architectural concept of augmenting transformer layers with an explicit, non-differentiable memory cache of past token states to extend model capacity without back-propagation.
- Book: AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models, Junfeng Fang et al. (2025). It details the core challenge and baseline methodologies of sequential model editing and preserving factual knowledge without catastrophic interference, motivating the self-updatable latent memory pool in MemoryLLM.
- Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). It establishes essential benchmark protocols and evaluation methodologies for sequential knowledge editing and factual updates in language models.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). It provides a key system design paradigm for treating LLM memory hierarchically to overcome context limits, establishing the functional context for internal latent memory mechanisms.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). It explores parameter-free core continual learning via external prompt memory pools, offering foundational principles for updating model behavior without full parameter retraining.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It outlines the foundational theoretical framework for reading and writing to explicit memory stores integrated end-to-end with neural networks.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It provides a comprehensive theoretical taxonomy of continual learning and stability-plasticity trade-offs that underpin the controlled forgetting dynamics of MemoryLLM.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). It extends the paradigm of dynamic model memory by training a dedicated, modular memory model to parametrically supply up-to-date knowledge to frozen executive LLMs.
- Paper: $δ$-mem: Efficient Online Memory for Large Language Models, Jingdi Lei et al. (2026). It advances lightweight online memory updates by integrating an evolving, low-parameter state into attention mechanisms during multi-turn interactions.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). It develops neural test-time memory architectures with gradient-based surprise metrics and adaptive forgetting gates to dynamically absorb long-range context.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). It expands continual test-time learning directly into standard transformer sub-layers to achieve scalable context compression and memorization during inference.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). It investigates multi-mechanism continual learning compositions for long-horizon memorization and retention under extensive sequential updates.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It provides a dedicated evaluation benchmark to assess how well dynamic memory and long-context systems retain temporal consistency and factual information over extended deployment periods.
