Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory

Xin YuZiyu XiaoZengxuan WenZelin LiJiaxi ZhouHualei WangHaohua WangHaizhen HuangWeiwei DengFeng Sun

article2026ACL11 citations

Proposes Mnemis, a long-term memory framework for large language models that pairs similarity search with deliberate top-down traversal over hierarchical graphs to resolve complex queries requiring global context.

Listen

As artificial intelligence systems transition into persistent, interactive agents, managing long-term conversational memory across extended interactions has become critical. Standard memory architectures—such as Retrieval-Augmented Generation and graph-based memory systems—rely almost exclusively on text matching and vector similarity. While fast and effective for straightforward lookup tasks, this fast, heuristic retrieval approach struggles when user requests require broad global reasoning, exhaustive evidence collection across time, or structured multi-hop deduction across months of conversation history.

The article introduces and evaluates Mnemis, a dual-route AI memory framework designed to overcome these retrieval limitations. The main objective of the article is to demonstrate how combining intuitive similarity search with structured, top-down exploration over hierarchical knowledge graphs significantly enhances an AI agent's long-term recall and reasoning capabilities.

To achieve this, Mnemis organizes conversational history across two connected storage layers: a base graph capturing granular entities, relationships, and raw text episodes, and a hierarchical graph abstracting entities into multi-level categories. The hierarchical graph is constructed using three governing principles: minimal concept abstraction to preserve detail, many-to-many mapping allowing nodes to belong to multiple categories across different contexts, and compression efficiency constraints to keep categories balanced. Memory retrieval uses a dual-route mechanism: a fast similarity search based on text and vector embeddings, paired with a deliberate global selection route that navigates the category hierarchy from top to bottom. The retrieved candidates from both routes are then unified and ranked by a secondary ranking model before being passed to the primary language model.

The authors evaluated Mnemis against numerous memory baselines across two long-term conversational benchmarks: LoCoMo, comprising approximately 1,540 evaluated questions across 16,000-token user histories, and LongMemEval-S, spanning 500 sessions averaging 115,000 tokens. Key findings indicate that Mnemis achieves top-tier performance across both benchmarks, scoring 93.9 on LoCoMo and 91.6 on LongMemEval-S when using GPT-4.1-mini, outperforming all compared baseline methods. Furthermore, the performance advantage was most pronounced on complex multi-hop and temporal reasoning tasks; for instance, on LoCoMo multi-hop questions, Mnemis scored 91.8 compared to 77.2 for full-context processing and 64.9 for standard retrieval. The experiments also revealed that relying on full model context windows degraded sharply as context grew to 115,000 tokens, whereas Mnemis maintained stable accuracy. Ablation studies confirmed that the performance gains stem directly from combining both retrieval routes, as neither route achieved optimal performance on its own.

These findings demonstrate that expanding native context windows or relying solely on similarity matching is insufficient and cost-ineffective for persistent agent deployments. By combining hierarchical graph traversal with similarity search, organizations can achieve higher response accuracy and reliability while maintaining compact context budgets. The modular architecture also allows organizations to deploy smaller, cost-effective ranking and embedding models without significant performance loss, enabling scalable long-term user personalization.

Organizations developing long-term AI agents should implement structured, multi-tier graph memory architectures rather than relying solely on raw context ingestion or simple vector search. Practical implementation should incorporate pruning rules such as early stopping during hierarchy navigation to manage computational costs. Future engineering efforts should focus on optimizing incremental hierarchy updates to avoid full periodic graph rebuilds, and exploring traversal mechanisms for multi-modal data formats beyond text.

The primary limitations noted in the article include the added operational complexity and processing overhead of maintaining and traversing hierarchical graphs, as well as the current requirement to periodically rebuild the hierarchy when base memories change. However, given that evaluations were conducted across standardized multi-session benchmarks and validated through both automated judging and human evaluation samples, decision-makers can place high confidence in Mnemis's demonstrated accuracy and architectural advantages for long-term memory management.

No sufficiently relevant recommendations were found.

Cover for Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory

Abstract

AI Memory, specifically how models organizes and retrieves historical messages, becomes increasingly valuable to Large Language Models (LLMs), yet existing methods (RAG and Graph-RAG) primarily retrieve memory through similarity-based mechanisms. While efficient, such System-1-style retrieval struggles with scenarios that require global reasoning or comprehensive coverage of all relevant information. In this work, We propose Mnemis, a novel memory framework that integrates System-1 similarity search with a complementary System-2 mechanism, termed Global Selection. Mnemis organizes memory into a base graph for similarity retrieval and a hierarchical graph that enables top-down, deliberate traversal over semantic hierarchies. By combining the complementary strength from both retrieval routes, Mnemis retrieves memory items that are both semantically and structurally relevant. Mnemis achieves state-of-the-art performance across all compared methods on long-term memory benchmarks, scoring 93.9 on LoCoMo and 91.6 on LongMemEval-S using GPT-4.1-mini.

Table of Contents

  • 1 Introduction
  • 2 Mnemis Methodology
  • 2.1 Base Graph
  • 2.2 Hierarchical Graph
  • 2.3 Memory Retrieval Mechanisms
  • 3 Experiments
  • 3.1 Experiment Setups
  • 3.2 Experiment Results
  • 3.3 Ablation Study
  • 3.3.1 Influence of System-1 and System-2 Routing
  • 3.3.2 Impact of top-kk
  • 3.3.3 Effect of Backend Models
  • 4 Related Work
  • 5 Conclusion
  • References
  • A Case Study
  • B Further Benchmark Results
  • B.1 Evaluation Metrics
  • C Detailed Performance Impact of top-kk
  • D Prompts

Knowls

  1. Knowl 1 — Mnemis combines similarity retrieval with hierarchical global selection

    model/method

    Mnemis is a long-term LLM memory framework with two linked stores and retrieval routes. Its base graph stores historical episodes together with extracted entities and relationship edges, supporting System-1 retrieval by textual and embedding similarity. Its hierarchical graph organizes entities into abstract categories, supporting System-2 retrieval by top-down selection through the hierarchy. For a user query, Mnemis gathers memory items from one or both routes, re-ranks episodes, entities or categories, and edges separately, then gives the resulting memory context and query to an answer-generating LLM. The design aims to combine fine-grained semantic matches with structurally relevant information that may be difficult to find through similarity alone.

  2. Knowl 2 — Base-graph representation and incremental memory ingestion

    model/method

    Mnemis represents memory in a base graph with four item types. An Episode is a raw historical text with a timestamp and an embedding. An Entity is a concrete person, organization, place, object, event, or well-defined concept; it has a name, contextual summary, type or role tag, and references to episodes in which it occurs. An Edge is a temporally or contextually scoped, verifiable fact or relationship involving specified entities; its fact text is embedded, and its validity interval is recorded with valid and invalid times. An Episodic Edge links an entity to episodes where it appears.

    Ingestion is incremental. New input is formatted as an Episode, and recent timestamped episodes provide context for extraction. An LLM identifies entity names, reflects to catch omissions, and de-duplicates candidates against existing entities using full-text and name-embedding similarity searches. It then extracts entity summaries, tags, and episode references. Edges are extracted from episode and entity context, followed by reflection and de-duplication. The speaker is also forcibly extracted as an entity.

  3. Knowl 3 — Hierarchical graph construction constraints

    model/method

    Mnemis builds a hierarchical graph bottom-up, treating base-graph entities as layer 0 and LLM-generated categories as higher-layer nodes. Category nodes carry entity-like descriptive fields plus their hierarchy layer; category edges connect them to child categories or entities. Categories are intended to capture shared meaning while remaining as specific as possible, and a lower-layer node may belong to multiple categories so that distinct semantic facets remain accessible.

    Construction applies two compression constraints. Each category must have at least the configured compression-ratio number nn of children; a node that cannot naturally be grouped may instead be promoted as a standalone category. Starting from layer 2, each layer must contain no more nodes than the layer beneath it; if this count-reduction rule is violated, construction terminates. At each layer, the LLM groups nodes using their names and tags, reuses applicable existing categories, and assigns nodes to categories. Construction stops when a compression constraint fails or a maximum layer limit is reached. The paper also describes a dedicated Speaker category for speaker and first-person references. The hierarchy is periodically rebuilt after base-graph updates; optimizing incremental hierarchy updates is left for future work.

  4. Knowl 4 — System-1 similarity search

    algorithm

    Given a user query and a base graph, System-1 retrieves episodes, entities, and edges separately using both embedding search and BM25 full-text search. For episodes, embedding search uses episode embeddings and BM25 searches episode content; for entities, the corresponding fields are name and summary; for edges, they are fact embeddings and fact text. Each search produces a ranked candidate list for its item type. Mnemis merges the embedding and BM25 rankings with reciprocal rank fusion, summing reciprocal candidate ranks, then orders candidates by the resulting fusion score. It performs this fusion separately for episodes, entities, and edges and truncates each list to its configured retrieval budget. In the main experiments, the budget was top-k=10k=10 episodes and top-2k=202k=20 entities and 20 edges.

  5. Knowl 5 — System-2 top-down global selection

    algorithm

    Given a query and Mnemis's hierarchical graph, System-2 begins at the top layer and asks an LLM to select all categories that may help answer the query, using category names and tags. It then presents the selected categories' children for another selection step, repeating layer by layer until reaching entities at layer 0. Selection is deliberately inclusive: nodes may be selected for direct relevance, useful context or personalization, or because they may contain helpful descendants. There is no fixed top-kk limit on category selection. At the lowest layer, Mnemis retrieves all episodes and edges directly connected to selected entities, plus entities connected through those edges. Traversal can stop early when the LLM judges that every child of a branch is relevant, avoiding further descent. The paper expects this route to be useful for enumerating relevant items; because the query remains unchanged during traversal, it may be less effective for reconstructing the order of temporal events.

  6. Knowl 6 — Benchmark and evaluation setup

    experimental setup

    Mnemis is evaluated on LoCoMo and LongMemEval-S using an LLM-as-a-Judge score with each benchmark's official judging prompt; scores are reported as 0/1 accuracy. LoCoMo contains roughly 2,000 questions over long conversations from 10 users, averaging about 600 turns, 32 sessions, and 16K tokens per user. Its reported overall scores exclude the adversarial question category, leaving 1,540 questions across multi-hop, temporal, open-domain, and single-hop categories. LongMemEval-S has 500 questions over sessions averaging about 115K tokens and tests single-session user information, multi-session reasoning, preferences, temporal reasoning, knowledge updates, and single-session assistant information.

    The main Mnemis configuration uses GPT-4.1-mini for memory processing and answering, Qwen3-Embedding-0.6B with 128-dimensional embeddings, and Qwen3-Reranker-8B. The answer context is limited to 10 episodes, 20 entities or categories, and 20 edges. Baseline comparisons include systems using GPT-4o-mini or GPT-4.1-mini where results were available; the paper also evaluates an episode-only RAG baseline with other settings aligned to Mnemis.

  7. Knowl 7 — LoCoMo performance

    empirical result

    On LoCoMo's 1,540 non-adversarial questions, Mnemis with GPT-4.1-mini scores 93.3 overall at the main setting of top-k=10k=10 episodes, 20 entities, and 20 edges. Increasing the episode budget to k=30k=30 yields 93.9 overall, the highest reported score in the compared results; EverMemOS is the strongest listed comparison at 92.3 with GPT-4.1-mini. For Mnemis at k=30k=30, the LLM-judge scores are 92.9 for multi-hop, 90.7 for temporal, 79.2 for open-domain, and 97.1 for single-hop questions. These results are reported under the paper's exclusion of LoCoMo's adversarial category.

  8. Knowl 8 — LongMemEval-S performance

    empirical result

    On all 500 LongMemEval-S questions, Mnemis with GPT-4.1-mini achieves an overall LLM-as-a-Judge score of 91.6. Its scores are 98.6 for single-session-user questions, 86.5 for multi-session reasoning, 100.0 for single-session preferences, 86.5 for temporal reasoning, 93.6 for knowledge updates, and 100.0 for single-session-assistant questions. The strongest listed comparison is EMem-G at 84.9 overall with GPT-4.1-mini, so Mnemis's overall score is 6.7 points higher in this reported comparison.

  9. Knowl 9 — Ablation shows complementary gains from the two retrieval routes

    empirical result

    In a LoCoMo ablation using GPT-4.1-mini and top-k=10k=10, System-1 episode-only RAG scores 73.8 overall, while System-1 graph retrieval using entities and edges scores 81.6. Combining episodes, entities, and edges in System-1 raises the score to 89.1; replacing reciprocal-rank fusion with the Qwen3-Reranker-8B gives the same 89.1 overall. System-2 alone scores 87.7, and combining System-1 with System-2 scores 93.3. In the combined setting, scores are 91.8 multi-hop, 90.3 temporal, 82.3 open-domain, and 96.2 single-hop. The authors interpret the results as evidence that the additional gain from the full system comes primarily from global selection rather than from the re-ranker alone, and that episodes complement the more compressed graph representation.

  10. Knowl 10 — Retrieval-budget effects on LoCoMo

    empirical result

    A LoCoMo study with GPT-4.1-mini varies the episode budget kk across 5, 10, 30, and 50; the corresponding entity and edge budgets are each 2k2k. Overall scores at those settings are: System-1 RAG, 64.3, 73.8, 82.7, and 85.6; System-1 Graph, 79.4, 81.6, 84.7, and 86.2; System-1 RAG plus Graph, 86.8, 89.1, 90.6, and 91.8; System-2 only, 87.3, 87.7, 87.9, and 88.1; and System-1 plus System-2, 92.2, 93.3, 93.9, and 93.4. Thus, episode-only RAG is especially sensitive to a small budget, whereas graph retrieval starts higher and changes less; the combined routes remain strongest across all tested budgets, peaking at k=30k=30. For RAG, the multi-hop score rises from 49.6 at k=5k=5 to 81.6 at k=50k=50, consistent with scattered evidence being missed under a narrow retrieval budget.

Coverage note — Omitted the secondary backend-model and re-ranker sensitivity analyses, per-stage token and runtime accounting, and illustrative case studies and prompt text; these are supporting detail rather than additional core methods or benchmark findings.

References

  1. 1.Muhammad Arslan, Hussam Ghanem, Saba Munawar, and Christophe Cruz. A survey on rag with llms. Procedia computer science, 246:3781–3790, 2024.
  2. 2.Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time. arXiv preprint arXiv:2501.00663, 2024.
  3. 3.Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024.
  4. 4.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory, 2025. https://arxiv.org/abs/2504.19413.
  5. 5.Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pages 758–759, 2009.
  6. 6.Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024.
  7. 7.Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, and Jiawei Han. Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025. https://arxiv.org/abs/2503.09516.
  8. 8.Daniel Kahneman. Thinking, fast and slow. macmillan, 2011.
  9. 9.Sangyeop Kim, Yohan Lee, Sanghwa Kim, Hyunjong Kim, and Sungzoon Cho. Pre-storage reasoning for episodic memory: Shifting inference burden to memory for personalized dialogue. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 22096–22113, 2025.
  10. 10.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474, 2020.
  11. 11.Zhiyu Li, Chenyang Xi, Chunyu Li, Ding Chen, Boyu Chen, Shichao Song, Simin Niu, Hanyu Wang, Jiawei Yang, Chen Tang, Qingchen Yu, Jihao Zhao, Yezhaohui Wang, Peng Liu, Zehao Lin, Pengyuan Wang, Jiahao Huo, Tianyi Chen, Kai Chen, Kehang Li, Zhen Tao, Huayi Lai, Hao Wu, Bo Tang, Zhengren Wang, Zhaoxin Fan, Ningyu Zhang, Linfeng Zhang, Junchi Yan, Mingchuan Yang, Tong Xu, Wei Xu, Huajun Chen, Haofen Wang, Hongkang Yang, Wentao Zhang, Zhi-Qin John Xu, Siheng Chen, and Feiyu Xiong. Memos: A memory os for ai system, 2025a. https://arxiv.org/abs/2507.03724.
  12. 12.Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, and Cheng-Lin Liu. Tl;dr: Too long, do re-weighting for efficient llm reasoning compression, 2025b. https://arxiv.org/abs/2506.02678.
  13. 13.Zhuowan Li, Cheng Li, Mingyang Zhang, Qiaozhu Mei, and Michael Bendersky. Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 881–893, 2024.
  14. 14.Xiao Liang, Zhongzhi Li, Yeyun Gong, Yelong Shen, Ying Nian Wu, Zhijiang Guo, and Weizhu Chen. Beyond pass@1: Self-play with variational problem synthesis sustains rlvr, 2025. https://arxiv.org/abs/2508.14029.
  15. 15.Zhenghao Lin, Zihao Tang, Xiao Liu, Yeyun Gong, Yi Cheng, Qi Chen, Hang Li, Ying Xin, Ziyue Yang, Kailai Yang, Yu Yan, Xiao Liang, Shuai Lu, Yiming Huang, Zheheng Luo, Lei Qu, Xuan Feng, Yaoxiang Wang, Yuqing Xia, Feiyang Chen, Yuting Jiang, Yasen Hu, Hao Ni, Binyang Li, Guoshuai Zhao, Jui-Hao Chiang, Zhongxin Guo, Chen Lin, Kun Kuang, Wenjie Li, Yelong Shen, Jian Jiao, Peng Cheng, and Mao Yang. Sigma: Differential rescaling of query, key and value for efficient language models, 2025. https://arxiv.org/abs/2501.13629.
  16. 16.Jiaheng Liu, Dawei Zhu, Zhiqi Bai, Yancheng He, Huanxuan Liao, Haoran Que, Zekun Wang, Chenchen Zhang, Ge Zhang, Jiebin Zhang, Yuanxing Zhang, Zhuo Chen, Hangyu Guo, Shilong Li, Ziqiang Liu, Yong Shan, Yifan Song, Jiayi Tian, Wenhao Wu, Zhejian Zhou, Ruijie Zhu, Junlan Feng, Yang Gao, Shizhu He, Zhoujun Li, Tianyu Liu, Fanyu Meng, Wenbo Su, Yingshui Tan, Zili Wang, Jian Yang, Wei Ye, Bo Zheng, Wangchunshu Zhou, Wenhao Huang, Sujian Li, and Zhaoxiang Zhang. A comprehensive survey on long context language modeling, 2025. https://arxiv.org/abs/2503.17407.
  17. 17.Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. Evaluating very long-term conversational memory of llm agents. arXiv preprint arXiv:2402.17753, 2024.
  18. 18.Jiayan Nan, Wenquan Ma, Wenlong Wu, and Yize Chen. Nemori: Self-organizing agent memory inspired by cognitive science, 2025. https://arxiv.org/abs/2508.03341.
  19. 19.Siru Ouyang, Jun Yan, I Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long T Le, Samira Daruki, Xiangru Tang, et al. Reasoningbank: Scaling agent self-evolving with reasoning memory. arXiv preprint arXiv:2509.25140, 2025.
  20. 20.Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo, Hao Cheng, Dongsheng Li, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Jianfeng Gao. Secom: On memory construction and retrieval for personalized conversational agents. In The Thirteenth International Conference on Learning Representations, 2025. https://openreview.net/forum?id=xKDZAW0He3.
  21. 21.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient context window extension of large language models. In The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=wHBfxhZu1u.
  22. 22.Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. Zep: A temporal knowledge graph architecture for agent memory, 2025. https://arxiv.org/abs/2501.13956.
  23. 23.Henrique* Schechter Vera, Sahil* Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiyi Wang, Zhe Li, Gus Martins, Jinhyuk Lee, Mark Sherwood, Juyeong Ji, Renjie Wu, Jingxiao Zheng, Jyotinder Singh, Abheesht Sharma, Divya Sreepat, Aashi Jain, Adham Elarabawy, AJ Co, Andreas Doumanoglou, Babak Samari, Ben Hora, Brian Potetz, Dahun Kim, Enrique Alfonseca, Fedor Moiseev, Feng Han, Frank Palma Gomez, Gustavo Hernández Ábrego, Hesen Zhang, Hui Hui, Jay Han, Karan Gill, Ke Chen, Koert Chen, Madhuri Shanbhogue, Michael Boratko, Paul Suganthan, Sai Meher Karthik Duddu, Sandeep Mariserla, Setareh Ariafar, Shanfeng Zhang, Shijie Zhang, Simon Baumgartner, Sonam Goenka, Steve Qiu, Tanmaya Dabral, Trevor Walker, Vikram Rao, Waleed Khawaja, Wenlei Zhou, Xiaoqi Ren, Ye Xia, Yichang Chen, Yi-Ting Chen, Zhe Dong, Zhongli Ding, Francesco Visin, Gaël Liu, Jiageng Zhang, Kathleen Kenealy, Michelle Casbon, Ravin Kumar, Thomas Mesnard, Zach Gleicher, Cormac Brick, Olivier Lacombe, Adam Roberts, Yunhsuan Sung, Raphael Hoffmann, Tris Warkentin, Armand Joulin, Tom Duerig, and Mojtaba Seyedhosseini. Embeddinggemma: Powerful and lightweight text representations. 2025. https://arxiv.org/abs/2509.20354.
  24. 24.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210–31227. PMLR, 2023.
  25. 25.Zihao Tang, Zheqi Lv, Shengyu Zhang, Fei Wu, and Kun Kuang. Modelgpt: Unleashing llm’s capabilities for tailored model generation, 2024. https://arxiv.org/abs/2402.12408.
  26. 26.Endel Tulving et al. Episodic and semantic memory. Organization of memory, 1(381-403):1, 1972.
  27. 27.Liang Wang, Haonan Chen, Nan Yang, Xiaolong Huang, Zhicheng Dou, and Furu Wei. Chain-of-retrieval augmented generation. In NeurIPS 2025, October 2025. https://www.microsoft.com/en-us/research/publication/chain-of-retrieval-augmented-generation/.
  28. 28.Yu Wang and Xi Chen. Mirix: Multi-agent memory system for llm-based agents, 2025. https://arxiv.org/abs/2507.07957.
  29. 29.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. Longmemeval: Benchmarking chat assistants on long-term interactive memory. 2024a. https://arxiv.org/abs/2410.10813.
  30. 30.Tao Wu, Mengze Li, Jingyuan Chen, Wei Ji, Wang Lin, Jinyang Gao, Kun Kuang, Zhou Zhao, and Fei Wu. Semantic alignment for multimodal large language models. In Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prabhakaran, Susanne Boll, Ramanathan Subramanian, Liang Zheng, Vivek K. Singh, Pablo César, Lexing Xie, and Dong Xu, editors, Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024, pages 3489–3498. ACM, 2024b. doi: 10.1145/3664647.3681014. https://doi.org/10.1145/3664647.3681014.
  31. 31.Tao Wu, Jingyuan Chen, Wang Lin, Mengze Li, Yumeng Zhu, Ang Li, Kun Kuang, and Fei Wu. Embracing imperfection: Simulating students with diverse cognitive levels using llm-based agents. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pages 9887–9908. Association for Computational Linguistics, 2025a. https://aclanthology.org/2025.acl-long.488/.
  32. 32.Tao Wu, Jingyuan Chen, Wang Lin, Jian Zhan, Mengze Li, Kun Kuang, and Fei Wu. Personalized distractor generation via mcts-guided reasoning reconstruction. CoRR, abs/2508.11184, 2025b. doi: 10.48550/ARXIV.2508.11184. https://doi.org/10.48550/arXiv.2508.11184.
  33. 33.Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Kristian Kersting, Jeff Z Pan, Hinrich Schütze, et al. Memory-r1: Enhancing large language model agents to manage and utilize memories via reinforcement learning. arXiv preprint arXiv:2508.19828, 2025.
  34. 34.Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176, 2025.
  35. 35.Sizhe Zhou. A simple yet strong baseline for long-term conversational memory of llm agents, 2025. https://arxiv.org/abs/2511.17208.

Citation

MLA
Tang, Z., et al. “Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory”. arXiv, 2026, http://arxiv.org/abs/2602.15313v2.
APA
Tang, Z., Yu, X., Xiao, Z., Wen, Z., Li, Z., Zhou, J., Wang, H., Wang, H., Huang, H., Deng, W., Sun, F., & Zhang, Q. (2026). Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory. arXiv. http://arxiv.org/abs/2602.15313v2
Chicago
Tang, Z., X. Yu, Z. Xiao, et al. 2026. “Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory”. arXiv. http://arxiv.org/abs/2602.15313v2.
Harvard
Tang, Z. et al. (2026) “Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.15313v2.
Vancouver
1. Tang Z, Yu X, Xiao Z, et al (2026) Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory. arXiv

BibTeX

@article{tang2026mnemis,
  title = {Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory},
  author = {Tang, Zihao and Yu, Xin and Xiao, Ziyu and Wen, Zengxuan and Li, Zelin and Zhou, Jiaxi and Wang, Hualei and Wang, Haohua and Huang, Haizhen and Deng, Weiwei and Sun, Feng and Zhang, Qi},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.15313v2},
  eprint = {2602.15313}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/