Human-Inspired Memory Architecture for LLM Agents

Doga KeresteciogluAlexei RobskyClemens VastersAnshul SharmaYitzhak Kesselman

article2026arXiv1 citations

Proposes a biologically grounded memory architecture for LLM agents that incorporates mechanisms such as sleep-phase consolidation and engram maturation to cut storage requirements by 58% while maintaining retrieval accuracy over long interaction horizons.

Listen

Large Language Model agents face a major limitation in enterprise environments: they lack persistent, adaptive memory across extended interactions. Current methods either discard context between sessions, rely on expanding prompt windows that increase computing costs without improving intelligence, or use simple vector databases that treat all information equally and cannot forget or consolidate data over time. The article presents and evaluates a biologically grounded memory architecture inspired by human cognition. The system maps three memory tiers—short-term cache, episodic vector storage, and a long-term semantic knowledge graph—and implements mechanisms including offline consolidation, adaptive forgetting, memory maturation, and reconsolidation upon retrieval.

To evaluate the system without benchmark data leakage, the researchers developed a synthetic calibration method that established all pipeline thresholds using independent, model-generated text. The architecture was tested across two benchmarks using a streaming protocol that processed events in strict chronological order. The first evaluation analyzed 13,127 software tracking issues comprising 120,000 events to measure retention precision. The second used a conversational benchmark across both small-scale settings (50 sessions) and large-scale streaming settings (475 sessions comprising roughly 540,000 turns) across multiple token budgets.

The findings show that deduplication-based consolidation significantly improves store efficiency and precision. In the software issue dataset, the pipeline achieved 97.2% retention precision with a 58% reduction in memory store size, outperforming the baseline by 21.8 percentage points. On large-scale conversational benchmarks at a 200,000-token budget, the pipeline matched raw retrieval accuracy within statistical parity (70.1% versus 71.2%) while showing slight improvements in multi-session reasoning (+1.2 percentage points) and temporal reasoning (+3.0 percentage points). Conversely, overly aggressive consolidation methods, such as clustering and summarization, degraded factual recall down to 48.4%, proving that memory consolidation should focus on deduplicating redundant data rather than aggressive summarization.

These results demonstrate that enterprise AI systems can reduce data storage overhead and operational costs without sacrificing recall performance. By establishing predictable memory decay and deduplication, organizations can maintain scalable agent systems that self-regulate memory size over long operational lifespans. The architecture provides a tunable operating curve, allowing organizations to balance context token costs against required task accuracy based on their specific application needs.

Decision-makers should consider adopting deduplication-focused memory pipelines for long-running operational agents, avoiding aggressive summarization that discards essential factual detail. Next steps should include testing these memory pipelines in end-to-end task environments, such as automated issue triage, and directly comparing performance against other agent frameworks. Readers should note that two mechanisms—memory maturation and reconsolidation—require real-world multi-week deployments with repeated queries and contradictions to fully demonstrate their individual benefits, as current benchmark structures do not fully exercise these biological features.

arXiv: 2605.08538
  • Paper: Why There Are Complementary Learning Systems in the Hippocampus and Neocortex, James L. McClelland et al. (1995). This foundational paper establishes the complementary learning systems theory of dual-memory hippocampal-neocortical consolidation that directly inspires the cognitive sleep-phase consolidation and engram maturation mechanisms of the source architecture.
  • Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). It introduces tiered, operating-system-inspired memory management for conversational LLM agents, serving as a critical architectural antecedent to cognitive agent memory pipelines.
  • Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It provides the foundational framework and baseline insights for evaluating very long-term conversational memory in LLM agents over multi-session horizons.
  • Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). It explores controlled forgetting and self-updatable memory pools in language models, directly informing the source's interference-based forgetting mechanisms.
  • Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It details the core principles of continual learning and the stability-plasticity tradeoff necessary for mitigating catastrophic forgetting in persistent agent memory.
  • Paper: Memory Networks, Jason Weston et al. (2014). It introduces the concept of augmenting neural networks with explicit, read-write external memory stores upon which modern LLM agent memory architectures are built.
Cover for Human-Inspired Memory Architecture for LLM Agents

Abstract

Current LLM agents lack principled mechanisms for managing persistent memory across long interaction horizons. We present a biologically-grounded memory architecture comprising six cognitive mechanisms: (1) sleep-phase consolidation, (2) interference-based forgetting, (3) engram maturation, (4) reconsolidation upon retrieval, (5) entity knowledge graphs, and (6) hybrid multi-cue retrieval. Each mechanism addresses a specific failure mode of naive memory accumulation. We introduce a synthetic calibration methodology that derives all pipeline thresholds without benchmark data exposure, eliminating a common source of evaluation leakage. We evaluate on two benchmarks. First, a VSCode issue-tracking dataset (13K issues, 120K events) where deduplication-based consolidation achieves 97.2% retention precision with 58% store reduction (+21.8 pp over baseline). Second, the LongMemEval personal-chat benchmark where we conduct the first streaming M-tier evaluation (475 sessions, ~540K unique turns). At a 200K-token context budget, our pipeline matches raw retrieval accuracy (70.1% vs. 71.2%, overlapping 95% CI) while exposing a tunable accuracy/store-size operating curve. At S-tier scale (50 sessions), dedup-based consolidation yields a +13.3 pp improvement in preference recall.

Table of Contents

  • 1 Introduction
  • 2 Biological Foundations
  • 3 Technical Architecture
  • 4 Memory Consolidation Pipeline
  • 5 Adaptive Forgetting
  • 6 Memory Maturation Dynamics
  • 7 Retrieval and Agent Integration
  • 7.1 Hybrid Retrieval
  • 7.2 Reconsolidation
  • 8 Experimental Methodology
  • 8.1 Synthetic Calibration
  • 8.2 Evaluation Protocol
  • 9 Evaluation
  • 9.1 VSCode Issue Tracking
  • 9.2 LongMemEval: S-Tier (50 Sessions)
  • 9.3 LongMemEval: M-tier (475 Sessions)
  • 10 Related Work
  • 11 Limitations
  • 12 Conclusion
  • References

Knowls

  1. Knowl 1 — Biologically Grounded Three-Tier Memory Architecture for LLM Agents

    model/method

    The agent memory architecture maps three biological memory tiers directly to computational storage subsystems:

    1. Short-Term Memory (Prefrontal Cortex →\to Hot Cache): In-memory storage with a time-to-live (TTL) of minutes to hours, dedicated to the immediate conversational session.
    2. Medium-Term Memory (Hippocampus →\to Warm Episodic Store): A time-indexed vector store preserving full-fidelity event records with a TTL of days to weeks.
    3. Long-Term Memory (Neocortex →\to Semantic Knowledge Graph): A permanent semantic graph capturing abstracted entities, relationships, and multi-hop associations.

    All three tiers share a unified underlying data layer, preventing inter-service data replication. Memory lifecycle operations (consolidation, forgetting, maturation, reconsolidation) are implemented as deterministic algorithmic procedures operating on embeddings and numerical scores without invoking large language models for background maintenance.

  2. Knowl 2 — Memory Consolidation and Importance Scoring Formulation

    model/method

    Memory consolidation operates as scheduled batch processing (defaulting to every 6 hours) mimicking biological sharp-wave ripples. Prior to score filtering, a temporal validation stage quarantines out-of-order, duplicate, or causally inverted events (quarantine TTL: 15 minutes).

    Each pending memory event ee is evaluated using a composite importance score:

    S(e)=∑i=15wi⋅fi(e)S(e) = \sum_{i=1}^5 w_i \cdot f_i(e)

    where fi(e)f_i(e) are scoring factors and wiw_i are their weights:

    • Recency (w1=0.25w_1 = 0.25): Exponential decay based on timestamp (hippocampal consolidation).
    • Frequency (w2=0.25w_2 = 0.25): Inverse frequency of similar events (Hebbian learning).
    • Bayesian Surprise (w3=0.20w_3 = 0.20): Embedding distance from prior event distribution (dopaminergic signaling).
    • Entity Salience (w4=0.15w_4 = 0.15): Maximum importance score among referenced entities (amygdala tagging).
    • Outcome (w5=0.15w_5 = 0.15): Goal completion and reward reinforcement signal.

    Events are partitioned by composite score: the top 20% are promoted to the permanent semantic graph as LLM-generated gists with entity links; the middle 60% are retained in the episodic store; and the bottom 20% are pruned.

  3. Knowl 3 — Adaptive Forgetting Dynamics and Progressive Graceful Degradation

    model/method

    The memory architecture implements two forgetting mechanisms and a multi-level degradation schedule:

    1. Passive Decay: For events awaiting consolidation, importance score I(t)I(t) decays exponentially:

    I(t)=I0⋅e−λtI(t) = I_0 \cdot e^{-\lambda t}

    where I0I_0 is the initial importance score, tt is the elapsed time in hours since encoding, and λ=0.001 h−1\lambda = 0.001\text{ h}^{-1} is the decay rate (half-life of approximately 29 days / 693 hours).

    1. Interference-Based Forgetting: When memories share overlapping semantic features or entities, an interference penalty is computed:

    Iinterference=∑jwj⋅sim(mi,mj)I_{\text{interference}} = \sum_j w_j \cdot \text{sim}(m_i, m_j)

    where sim(mi,mj)\text{sim}(m_i, m_j) denotes cosine similarity between memory embeddings, and wjw_j weights interference directions (wretroactive=0.6w_{\text{retroactive}} = 0.6, wproactive=0.4w_{\text{proactive}} = 0.4). Memories exhibiting high interference and low importance scores are selectively pruned.

    1. Graceful Degradation: Before complete deletion, memories step down through six discrete fidelity tiers: L0 (full episodic record, 100% fidelity), L2 (summary, 50%), L3 (gist, 25%), and L5 (tombstone record, 0% fidelity, retaining only metadata asserting that the event occurred). Degradation is governed by memory age combined with importance score.
  4. Knowl 4 — Engram Maturation Activation Dynamics

    equation

    When an event is consolidated into the semantic knowledge graph, the full-fidelity episodic trace remains immediately accessible in the episodic store, while the semantic graph entry is initialized with activation_strength=0.0\text{activation\_strength} = 0.0. The explicit retrieval activation strength A(t)A(t) evolves over time according to a sigmoid function:

    A(t)=11+e−(t−t1/2)/kA(t) = \frac{1}{1 + e^{-(t - t_{1/2})/k}}

    where:

    • t≥0t \ge 0 is the time elapsed since consolidation in hours,
    • t1/2=168 hourst_{1/2} = 168\text{ hours} (1 week) is the maturation half-life,
    • k=48 hoursk = 48\text{ hours} is the logistic slope parameter.

    At consolidation (t=0t = 0), the memory is silent (A≈0.03A \approx 0.03). At t=168 hourst = 168\text{ hours}, A(t)=0.5A(t) = 0.5 (the retrieval threshold). At t=336 hourst = 336\text{ hours} (2 weeks), the memory reaches full maturity (A>0.9A > 0.9). While A(t)<0.5A(t) < 0.5, memories cannot be explicitly retrieved but exert implicit priming effects that bias the relevance scoring of related active memories.

  5. Knowl 5 — Hybrid Multi-Cue Retrieval and Reconsolidation Mechanism

    model/method

    Retrieval queries execute across all three memory tiers simultaneously with prioritized merging:

    1. Short-Term Hot Cache: Scanned first for immediate in-session context (highest priority, sub-second latency).
    2. Warm Episodic Store: Vector similarity search over recent full-fidelity memories filtered by dynamic importance score thresholds.
    3. Semantic Knowledge Graph (GraphRAG): Multi-hop graph traversal over mature memories satisfying A(t)≥0.5A(t) \ge 0.5, seeded by top vector similarity matches.

    Candidate items across tiers are merged, deduplicated, and ranked with a recency boost.

    Reconsolidation: Retrieved memories enter a modifiable lability window (default duration: 60 minutes). If new conversational inputs introduce contradictions, updates, or elaborations during this window, the existing memory is updated in place via adaptive blending weighted by source confidence, recency, and contradiction severity. Decision outcomes reinforce or penalize the importance scores of retrieved memories.

  6. Knowl 6 — Zero-Leakage Synthetic Calibration Methodology

    experimental setup

    To prevent benchmark contamination and threshold tuning leakage, all pipeline parameters and decision thresholds are calibrated exclusively on synthetic LLM-generated corpora produced from fixed specifications with zero exposure to downstream evaluation benchmarks.

    • Similarity Thresholds: Derived from 8 topically diverse synthetic personal-chat sessions (88 turns) embedded with text-embedding-3-large (3072 dimensions):

      • Near-deduplication threshold: Set at the 99th percentile of all-pairs similarity (0.5590.559).
      • Cluster distance: Set to 1−P951 - P_{95} of within-session similarity (0.4040.404).
      • Interference threshold: Set to P90P_{90} of within-session similarity (0.5420.542).
    • Importance Weights: Derived from 50 synthetic multi-topic sessions (483 turns: 377 substantive, 106 filler across 14 topics over 3 simulated months) via ROC AUC analysis with AUC-excess normalization. Discriminative weights are: content length (AUC=0.77AUC = 0.77, weight 0.3630.363), turn position within session (weight 0.3250.325), Bayesian surprise (content length and embedding surprise), and recency (AUC=0.51AUC = 0.51, weight 0.0190.019).

  7. Knowl 7 — Retention Precision on VSCode Issue-Tracking Stream

    empirical result

    The memory consolidation and forgetting pipeline was evaluated on a temporal stream of 13,127 real VSCode GitHub issues comprising 120,000 events (issue creation, comments, label changes, state transitions, and assignments) between December 2025 and February 2026, embedded with text-embedding-3-large.

    • Retention Precision: Consolidation and forgetting achieved 97.2%97.2\% retention precision (predicting whether retained events are referenced in future issue activities) with a 58%58\% reduction in total store size, compared to 75.4%75.4\% retention precision for the keep-everything baseline (+21.8+21.8 percentage points improvement).
    • Store Self-Regulation: The episodic memory store stabilized at a constant footprint of 300–500300\text{--}500 events regardless of increasing input volume.
    • Optimal Decay Rate: The empirical optimal decay rate was λ=0.001 h−1\lambda = 0.001\text{ h}^{-1} (half-life ≈29\approx 29 days), indicating that operational software engineering agents require memory retention horizons tuned to domain cadence rather than circadian 24-hour cycles.
  8. Knowl 8 — LongMemEval S-Tier Benchmark Ablation Results

    data/table

    The table presents the 9-configuration ablation study on the LongMemEval S-tier benchmark (500 questions, ~50 sessions / ~500 turns per question) evaluated using GPT-4o with bootstrapped 95% confidence intervals (10,000 resamples). Tasks include Knowledge Update (KU), Multi-Session Reasoning (MS), Single-Session Preference (SS-P), Single-Session Assistant (SS-A), Single-Session User (SS-U), and Temporal Reasoning (Temp).

    Config KU MS SS-P SS-A SS-U Temp Overall (95% CI)
    Raw RAG (baseline) 84.6 67.7 56.7 94.6 95.7 74.4 78.4 [74.8, 82.0]
    Dedup-only 83.3 61.7 70.0 96.4 90.0 74.4 76.8 [73.0, 80.4]
    Dedup-adaptive-25K 82.1 62.4 66.7 98.2 90.0 74.4 76.8 [73.2, 80.4]
    Dedup-adaptive-50K 83.3 62.4 66.7 98.2 88.6 72.2 76.2 [72.6, 79.8]
    Dedup + recon 80.8 62.4 70.0 98.2 88.6 71.4 75.8 [72.0, 79.6]
    Dedup + hybrid 82.1 63.2 60.0 92.9 88.6 70.7 74.8 [71.0, 78.6]
    Adaptive-10K 71.8 29.3 50.0 92.9 62.9 59.4 57.0 [52.8, 61.4]
    Aggressive consol. 60.3 40.6 56.7 39.3 60.0 45.1 48.4 [44.0, 52.8]
    Aggressive + recon 59.0 37.6 66.7 44.6 60.0 45.9 48.8 [44.4, 53.4]

    Moderate deduplication configurations overlap the Raw RAG baseline 95% CI ([74.8, 82.0]), confirming non-destructive compression. In addition, deduplication provides a directional +13.3+13.3 percentage point boost in preference recall (SS-P: 56.7%→70.0%56.7\% \to 70.0\%). Aggressive consolidation (agglomerative clustering into summaries) drops accuracy to 48.4%48.4\%, proving that consolidation must perform exact deduplication rather than lossy summarization.

  9. Knowl 9 — LongMemEval Streaming M-Tier Evaluation

    data/table

    The table reports performance on the LongMemEval M-tier under temporal streaming conditions (475 sessions per question, ~4,900 turns per question, ~540,000 unique turns total) across adaptive token memory budgets using GPT-4o as judge, with 95% bootstrap confidence intervals (10,000 resamples).

    Config (Tokens) KU MS SS-P SS-A SS-U Temp Overall (95% CI)
    Raw RAG (k=10k = 10) 78.2 57.1 53.3 96.4 92.9 63.2 71.2 [67.2, 75.0]
    Dedup-adaptive 200K 74.4 58.3 53.3 91.1 85.7 66.2 70.1 [66.0, 74.2]
    Dedup-adaptive 115K 69.2 57.9 40.0 83.9 82.9 60.2 65.6 [61.4, 69.6]
    Dedup-adaptive 50K 56.4 39.1 43.3 60.7 50.0 51.1 49.2 [44.8, 53.6]
    Dedup-adaptive 25K 48.7 24.8 33.3 41.1 34.3 49.6 38.8 [34.6, 43.2]

    At a 200K-token context budget, the pipeline matches the Raw RAG baseline (70.1%70.1\% vs. 71.2%71.2\%, overlapping 95% CIs) while outperforming Raw RAG on multi-session reasoning (58.3%58.3\% vs. 57.1%57.1\%) and temporal reasoning (66.2%66.2\% vs. 63.2%63.2\%). The four budget tiers establish a controllable accuracy/memory-size operating curve, with the non-destructive crossover point occurring between 115K and 200K tokens.

  10. Knowl 10 — Benchmark Structural Constraints on Maturation and Reconsolidation Validation

    limitation

    Two of the six core architectural mechanisms could not be empirically isolated by ablation due to structural limitations of standard memory benchmarks:

    1. Engram Maturation: Maturation requires repeated retrieval over extended multi-week horizons to differentiate activation strengths. In LongMemEval, each test question operates on an isolated context haystack with no shared retrieval history across queries, preventing the accumulation of access frequencies. Consequently, evaluations ran with static uniform activation strength (A=1.0A = 1.0).
    2. Reconsolidation: Reconsolidation activates when incoming memories contradict prior knowledge. LongMemEval's synthetic personal-chat histories are constructed to be strictly non-conflicting by design. As a result, the Dedup+Recon configuration (75.8%75.8\% [72.0, 79.6]) exhibited complete CI overlap with Dedup-Only (76.8%76.8\% [73.0, 80.4]).

    Both mechanisms are retained in the architecture based on neurobiological theory and operational multi-user agent requirements rather than benchmark ablation evidence.

Coverage note — None omitted; all contributed architectural components, mathematical formulations, calibration procedures, empirical benchmark evaluations, and limitations are fully represented.

References

  1. 1.Michael C. Anderson. Rethinking interference theory: Executive control and the mechanisms of forgetting. Journal of Memory and Language, 49(4):415–445, 2003.
  2. 2.Paul W. Frankland and Bruno Bontempi. The organization of recent and remote memories. Nature Reviews Neuroscience, 6(2):119–130, 2005.
  3. 3.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526, 2017.
  4. 4.Takashi Kitamura, Sachie K. Ogawa, Dheeraj S. Roy, Teruhiro Okuyama, Mark D. Morrissey, Lillian M. Smith, Roger L. Redondo, and Susumu Tonegawa. Engrams and circuits crucial for systems consolidation of a memory. Science, 356(6333):73–78, 2017.
  5. 5.James L. McClelland, Bruce L. McNaughton, and Randall C. O’Reilly. Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. Psychological Review, 102(3):419–457, 1995.
  6. 6.Karim Nader, Glenn E. Schafe, and Joseph E. LeDoux. Fear memories require protein synthesis in the amygdala for reconsolidation after retrieval. Nature, 406(6797):722–726, 2000.
  7. 7.Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems. In Advances in Neural Information Processing Systems, volume 37, 2024.
  8. 8.Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023.
  9. 9.Ruobing Shao, Canwen Chen, Jianfei Jia, and Bo Xiao. RAISE: Retrieval-augmented interaction simulation engine for automatic llm agent evaluation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  10. 10.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems, volume 36, 2023.
  11. 11.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. LongMemEval: Benchmarking chat assistants on long-term interactive memory. In Proceedings of the International Conference on Learning Representations (ICLR), 2025.

Citation

MLA
Kerestecioglu, D., et al. “Human-Inspired Memory Architecture for LLM Agents”. arXiv, 2026, http://arxiv.org/abs/2605.08538v1.
APA
Kerestecioglu, D., Robsky, A., Vasters, C., Sharma, A., & Kesselman, Y. (2026). Human-Inspired Memory Architecture for LLM Agents. arXiv. http://arxiv.org/abs/2605.08538v1
Chicago
Kerestecioglu, D., A. Robsky, C. Vasters, A. Sharma, and Y. Kesselman. 2026. “Human-Inspired Memory Architecture for LLM Agents”. arXiv. http://arxiv.org/abs/2605.08538v1.
Harvard
Kerestecioglu, D. et al. (2026) “Human-Inspired Memory Architecture for LLM Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.08538v1.
Vancouver
1. Kerestecioglu D, Robsky A, Vasters C, Sharma A, Kesselman Y (2026) Human-Inspired Memory Architecture for LLM Agents. arXiv

BibTeX

@article{kerestecioglu2026human,
  title = {Human-Inspired Memory Architecture for LLM Agents},
  author = {Kerestecioglu, Doga and Robsky, Alexei and Vasters, Clemens and Sharma, Anshul and Kesselman, Yitzhak},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.08538v1},
  eprint = {2605.08538}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/