Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
Xin YuZiyu XiaoZengxuan WenZelin LiJiaxi ZhouHualei WangHaohua WangHaizhen HuangWeiwei DengFeng Sun
Proposes Mnemis, a long-term memory framework for large language models that pairs similarity search with deliberate top-down traversal over hierarchical graphs to resolve complex queries requiring global context.
As artificial intelligence systems transition into persistent, interactive agents, managing long-term conversational memory across extended interactions has become critical. Standard memory architectures—such as Retrieval-Augmented Generation and graph-based memory systems—rely almost exclusively on text matching and vector similarity. While fast and effective for straightforward lookup tasks, this fast, heuristic retrieval approach struggles when user requests require broad global reasoning, exhaustive evidence collection across time, or structured multi-hop deduction across months of conversation history.
The article introduces and evaluates Mnemis, a dual-route AI memory framework designed to overcome these retrieval limitations. The main objective of the article is to demonstrate how combining intuitive similarity search with structured, top-down exploration over hierarchical knowledge graphs significantly enhances an AI agent's long-term recall and reasoning capabilities.
To achieve this, Mnemis organizes conversational history across two connected storage layers: a base graph capturing granular entities, relationships, and raw text episodes, and a hierarchical graph abstracting entities into multi-level categories. The hierarchical graph is constructed using three governing principles: minimal concept abstraction to preserve detail, many-to-many mapping allowing nodes to belong to multiple categories across different contexts, and compression efficiency constraints to keep categories balanced. Memory retrieval uses a dual-route mechanism: a fast similarity search based on text and vector embeddings, paired with a deliberate global selection route that navigates the category hierarchy from top to bottom. The retrieved candidates from both routes are then unified and ranked by a secondary ranking model before being passed to the primary language model.
The authors evaluated Mnemis against numerous memory baselines across two long-term conversational benchmarks: LoCoMo, comprising approximately 1,540 evaluated questions across 16,000-token user histories, and LongMemEval-S, spanning 500 sessions averaging 115,000 tokens. Key findings indicate that Mnemis achieves top-tier performance across both benchmarks, scoring 93.9 on LoCoMo and 91.6 on LongMemEval-S when using GPT-4.1-mini, outperforming all compared baseline methods. Furthermore, the performance advantage was most pronounced on complex multi-hop and temporal reasoning tasks; for instance, on LoCoMo multi-hop questions, Mnemis scored 91.8 compared to 77.2 for full-context processing and 64.9 for standard retrieval. The experiments also revealed that relying on full model context windows degraded sharply as context grew to 115,000 tokens, whereas Mnemis maintained stable accuracy. Ablation studies confirmed that the performance gains stem directly from combining both retrieval routes, as neither route achieved optimal performance on its own.
These findings demonstrate that expanding native context windows or relying solely on similarity matching is insufficient and cost-ineffective for persistent agent deployments. By combining hierarchical graph traversal with similarity search, organizations can achieve higher response accuracy and reliability while maintaining compact context budgets. The modular architecture also allows organizations to deploy smaller, cost-effective ranking and embedding models without significant performance loss, enabling scalable long-term user personalization.
Organizations developing long-term AI agents should implement structured, multi-tier graph memory architectures rather than relying solely on raw context ingestion or simple vector search. Practical implementation should incorporate pruning rules such as early stopping during hierarchy navigation to manage computational costs. Future engineering efforts should focus on optimizing incremental hierarchy updates to avoid full periodic graph rebuilds, and exploring traversal mechanisms for multi-modal data formats beyond text.
The primary limitations noted in the article include the added operational complexity and processing overhead of maintaining and traversing hierarchical graphs, as well as the current requirement to periodically rebuild the hierarchy when base memories change. However, given that evaluations were conducted across standardized multi-session benchmarks and validated through both automated judging and human evaluation samples, decision-makers can place high confidence in Mnemis's demonstrated accuracy and architectural advantages for long-term memory management.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). GraphRAG’s hierarchical entity communities and global-query summaries provide the clearest foundation for understanding Mnemis’s hierarchical graph route beyond similarity search.
No sufficiently relevant recommendations were found.
