Human-Inspired Memory Architecture for LLM Agents
Doga KeresteciogluAlexei RobskyClemens VastersAnshul SharmaYitzhak Kesselman
Proposes a biologically grounded memory architecture for LLM agents that incorporates mechanisms such as sleep-phase consolidation and engram maturation to cut storage requirements by 58% while maintaining retrieval accuracy over long interaction horizons.
Large Language Model agents face a major limitation in enterprise environments: they lack persistent, adaptive memory across extended interactions. Current methods either discard context between sessions, rely on expanding prompt windows that increase computing costs without improving intelligence, or use simple vector databases that treat all information equally and cannot forget or consolidate data over time. The article presents and evaluates a biologically grounded memory architecture inspired by human cognition. The system maps three memory tiers—short-term cache, episodic vector storage, and a long-term semantic knowledge graph—and implements mechanisms including offline consolidation, adaptive forgetting, memory maturation, and reconsolidation upon retrieval.
To evaluate the system without benchmark data leakage, the researchers developed a synthetic calibration method that established all pipeline thresholds using independent, model-generated text. The architecture was tested across two benchmarks using a streaming protocol that processed events in strict chronological order. The first evaluation analyzed 13,127 software tracking issues comprising 120,000 events to measure retention precision. The second used a conversational benchmark across both small-scale settings (50 sessions) and large-scale streaming settings (475 sessions comprising roughly 540,000 turns) across multiple token budgets.
The findings show that deduplication-based consolidation significantly improves store efficiency and precision. In the software issue dataset, the pipeline achieved 97.2% retention precision with a 58% reduction in memory store size, outperforming the baseline by 21.8 percentage points. On large-scale conversational benchmarks at a 200,000-token budget, the pipeline matched raw retrieval accuracy within statistical parity (70.1% versus 71.2%) while showing slight improvements in multi-session reasoning (+1.2 percentage points) and temporal reasoning (+3.0 percentage points). Conversely, overly aggressive consolidation methods, such as clustering and summarization, degraded factual recall down to 48.4%, proving that memory consolidation should focus on deduplicating redundant data rather than aggressive summarization.
These results demonstrate that enterprise AI systems can reduce data storage overhead and operational costs without sacrificing recall performance. By establishing predictable memory decay and deduplication, organizations can maintain scalable agent systems that self-regulate memory size over long operational lifespans. The architecture provides a tunable operating curve, allowing organizations to balance context token costs against required task accuracy based on their specific application needs.
Decision-makers should consider adopting deduplication-focused memory pipelines for long-running operational agents, avoiding aggressive summarization that discards essential factual detail. Next steps should include testing these memory pipelines in end-to-end task environments, such as automated issue triage, and directly comparing performance against other agent frameworks. Readers should note that two mechanisms—memory maturation and reconsolidation—require real-world multi-week deployments with repeated queries and contradictions to fully demonstrate their individual benefits, as current benchmark structures do not fully exercise these biological features.
- Paper: Why There Are Complementary Learning Systems in the Hippocampus and Neocortex, James L. McClelland et al. (1995). This foundational paper establishes the complementary learning systems theory of dual-memory hippocampal-neocortical consolidation that directly inspires the cognitive sleep-phase consolidation and engram maturation mechanisms of the source architecture.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). It introduces tiered, operating-system-inspired memory management for conversational LLM agents, serving as a critical architectural antecedent to cognitive agent memory pipelines.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It provides the foundational framework and baseline insights for evaluating very long-term conversational memory in LLM agents over multi-session horizons.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). It explores controlled forgetting and self-updatable memory pools in language models, directly informing the source's interference-based forgetting mechanisms.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It details the core principles of continual learning and the stability-plasticity tradeoff necessary for mitigating catastrophic forgetting in persistent agent memory.
- Paper: Memory Networks, Jason Weston et al. (2014). It introduces the concept of augmenting neural networks with explicit, read-write external memory stores upon which modern LLM agent memory architectures are built.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). It advances the retrieval paradigm by shifting from static retrieval to active, graph-based memory reconstruction guided directly by multi-step LLM reasoning.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). It introduces an execution environment and reward model to benchmark and evaluate long-horizon memory dynamics across complex interactive agent tasks.
- Paper: Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents, Theo Rusu et al. (2026). It investigates graph-based memory representations and explicit selective forgetting mechanisms across persistent long-term dialogue benchmarks.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). It explores a parametric approach to external agent memory by synthesizing structured reflections and training a dedicated memory model to serve an executive LLM.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). It extends long-term memory architectures to multi-agent workflow automation by structuring and sharing modular procedural execution logs.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). It evaluates how well agents maintain state and memory when user intents dynamically evolve and contradict previous instructions across multi-turn interactions.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). It provides a comprehensive survey evaluating the holistic efficiency tradeoffs across memory, tool use, and planning in autonomous agent systems.
