In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents
Zhen TanJun YanI-Hung HsuRujun HanZifeng WangLong T. LeYiwen SongYanfei ChenHamid PalangiGeorge Lee
Proposes Reflective Memory Management, a framework that couples dynamic topic-based conversation summarization with online reinforcement learning from response attributions to overcome rigid memory structures and fixed retrievers in long-term dialogue agents.
Large Language Models often struggle to maintain coherence, personalization, and factual recall during long-term interactions across multiple sessions. In typical applications such as virtual assistants or customer service, existing external memory systems fail because they segment conversations using rigid, arbitrary boundaries like fixed turn or session limits and rely on static search mechanisms that cannot adapt to evolving dialogue contexts.
The article demonstrates that an adaptive memory framework combining forward-looking topic structuring with backward-looking retrieval tuning can significantly improve personalized dialogue performance. To achieve this, the article evaluates a new framework named Reflective Memory Management on two standardized multi-session conversational benchmarks, measuring retrieval accuracy, lexical similarity, and response correctness against existing retrieval and memory baselines.
The evaluated approach operates through two core components: Prospective Reflection and Retrospective Reflection. Prospective Reflection extracts conversation topics at the end of each session and dynamically integrates them into an external memory bank by either creating new topic entries or merging updates into existing ones. Retrospective Reflection pairs an initial search retriever with a lightweight reranking layer. During response generation, the underlying language model automatically generates inline citations indicating which retrieved memories were actually useful; these citations serve as unsupervised reinforcement learning rewards to iteratively train the reranker without requiring expensive labeled data.
The analysis reveals several key findings. First, the Reflective Memory Management framework consistently outperformed all evaluated baselines across benchmarks, achieving a 10% to 13% improvement in accuracy and recall on the LongMemEval benchmark compared to baseline retrieval methods without memory management. Second, topic-based organization proved superior to fixed turn or session boundaries, closely approaching the retrieval performance of an ideal oracle setup. Third, reinforcement learning based on attribution citations proved highly accurate, yielding useful memory identification scores of roughly 86% to 90% across precision, recall, and overall balance. Finally, testing showed that simply providing a very large context window was insufficient, as models operating solely on extended raw context failed to recall historical details accurately.
These results indicate that conversational systems do not require complex, white-box architectural changes or costly manual data labeling to maintain long-term personalization. Implementing an adaptive topic-merging layer and an attribution-driven reranker allows existing off-the-shelf language models to handle evolving user preferences efficiently while reducing irrelevant context distractions. Interestingly, smaller generator models such as Gemini-1.5-Flash outperformed larger models like Gemini-1.5-Pro within this framework, primarily because stricter alignment in larger models caused them to abstain more frequently when handling personal information.
Organizations developing personalized conversational agents should consider adopting topic-decomposed memory architectures and citation-based reranking mechanisms rather than relying solely on expanded context windows. If labeled data is already available, teams can implement offline pretraining before running online reinforcement learning updates to further accelerate performance gains. Developers should also carefully tune the number of retrieved and reranked memory entries (such as retrieving 50 and passing 10 to the model) to balance processing efficiency against retrieval accuracy.
The findings are subject to several boundary conditions. The current framework was evaluated exclusively on text-based interactions, meaning performance on multi-modal conversations involving audio or images remains unverified. In addition, reinforcement learning updates introduce computational overhead that could affect ultra-low-latency, real-time deployments. Deploying memory systems that store personal information across sessions also necessitates robust data encryption and privacy protections to ensure sensitive user data is handled securely.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT establishes the hierarchical external-memory approach that helps clarify how this work shifts from manual memory paging to adaptive topic organization and retrieval tuning.
- Paper: Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models, Harshita Chopra et al. (2026). Thinking Ahead extends the paper’s prospective-memory idea by using imagined future actions as adaptive search probes for retrieving distant user history.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). Human-Inspired Memory Architecture carries adaptive long-term memory further with consolidation, forgetting, and reconsolidation across multiple memory tiers.
- Paper: Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory, Xin Yu et al. (2026). Mnemis continues the search for stronger long-term recall by combining ordinary similarity retrieval with top-down exploration of hierarchical memory graphs.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). Memory is Reconstructed, Not Retrieved extends adaptive memory access into iterative graph traversal, where each reasoning step guides what evidence to retrieve next.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). General Agentic Memory broadens adaptive memory beyond dialogue by pairing stored histories with iterative, query-directed research across varied long-context tasks.
- Paper: $δ$-mem: Efficient Online Memory for Large Language Models, Jingdi Lei et al. (2026). δ-mem explores a complementary continuation: encoding interaction history into a compact online state rather than relying on an external topic-organized memory bank.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). MemGym extends memory-system evaluation from static conversational recall to active decisions about what agents retain, summarize, and discard during long-horizon tasks.
