In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents

Zhen TanJun YanI-Hung HsuRujun HanZifeng WangLong T. LeYiwen SongYanfei ChenHamid PalangiGeorge Lee

article2025ACL130 citations

Proposes Reflective Memory Management, a framework that couples dynamic topic-based conversation summarization with online reinforcement learning from response attributions to overcome rigid memory structures and fixed retrievers in long-term dialogue agents.

Listen

Large Language Models often struggle to maintain coherence, personalization, and factual recall during long-term interactions across multiple sessions. In typical applications such as virtual assistants or customer service, existing external memory systems fail because they segment conversations using rigid, arbitrary boundaries like fixed turn or session limits and rely on static search mechanisms that cannot adapt to evolving dialogue contexts.

The article demonstrates that an adaptive memory framework combining forward-looking topic structuring with backward-looking retrieval tuning can significantly improve personalized dialogue performance. To achieve this, the article evaluates a new framework named Reflective Memory Management on two standardized multi-session conversational benchmarks, measuring retrieval accuracy, lexical similarity, and response correctness against existing retrieval and memory baselines.

The evaluated approach operates through two core components: Prospective Reflection and Retrospective Reflection. Prospective Reflection extracts conversation topics at the end of each session and dynamically integrates them into an external memory bank by either creating new topic entries or merging updates into existing ones. Retrospective Reflection pairs an initial search retriever with a lightweight reranking layer. During response generation, the underlying language model automatically generates inline citations indicating which retrieved memories were actually useful; these citations serve as unsupervised reinforcement learning rewards to iteratively train the reranker without requiring expensive labeled data.

The analysis reveals several key findings. First, the Reflective Memory Management framework consistently outperformed all evaluated baselines across benchmarks, achieving a 10% to 13% improvement in accuracy and recall on the LongMemEval benchmark compared to baseline retrieval methods without memory management. Second, topic-based organization proved superior to fixed turn or session boundaries, closely approaching the retrieval performance of an ideal oracle setup. Third, reinforcement learning based on attribution citations proved highly accurate, yielding useful memory identification scores of roughly 86% to 90% across precision, recall, and overall balance. Finally, testing showed that simply providing a very large context window was insufficient, as models operating solely on extended raw context failed to recall historical details accurately.

These results indicate that conversational systems do not require complex, white-box architectural changes or costly manual data labeling to maintain long-term personalization. Implementing an adaptive topic-merging layer and an attribution-driven reranker allows existing off-the-shelf language models to handle evolving user preferences efficiently while reducing irrelevant context distractions. Interestingly, smaller generator models such as Gemini-1.5-Flash outperformed larger models like Gemini-1.5-Pro within this framework, primarily because stricter alignment in larger models caused them to abstain more frequently when handling personal information.

Organizations developing personalized conversational agents should consider adopting topic-decomposed memory architectures and citation-based reranking mechanisms rather than relying solely on expanded context windows. If labeled data is already available, teams can implement offline pretraining before running online reinforcement learning updates to further accelerate performance gains. Developers should also carefully tune the number of retrieved and reranked memory entries (such as retrieving 50 and passing 10 to the model) to balance processing efficiency against retrieval accuracy.

The findings are subject to several boundary conditions. The current framework was evaluated exclusively on text-based interactions, meaning performance on multi-modal conversations involving audio or images remains unverified. In addition, reinforcement learning updates introduce computational overhead that could affect ultra-low-latency, real-time deployments. Deploying memory systems that store personal information across sessions also necessitates robust data encryption and privacy protections to ensure sensitive user data is handled securely.

arXiv: 2503.08026
  • Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT establishes the hierarchical external-memory approach that helps clarify how this work shifts from manual memory paging to adaptive topic organization and retrieval tuning.
Cover for In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents

Abstract

Large Language Models (LLMs) have made significant progress in open-ended dialogue, yet their inability to retain and retrieve relevant information from long-term interactions limits their effectiveness in applications requiring sustained personalization. External memory mechanisms have been proposed to address this limitation, enabling LLMs to maintain conversational continuity. However, existing approaches struggle with two key challenges. First, rigid memory granularity fails to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. Second, fixed retrieval mechanisms cannot adapt to diverse dialogue contexts and user interaction patterns. In this work, we propose Reflective Memory Management (RMM), a novel mechanism for long-term dialogue agents, integrating forward- and backward-looking reflections: (1) Prospective Reflection, which dynamically summarizes interactions across granularities—utterances, turns, and sessions—into a personalized memory bank for effective future retrieval, and (2) Retrospective Reflection, which iteratively refines the retrieval in an online reinforcement learning (RL) manner based on LLMs’ cited evidence. Experiments show that RMM demonstrates consistent improvement across various metrics and benchmarks. For example, RMM shows more than 10% accuracy improvement over the baseline without memory management on the LongMemEval dataset.

Knowls

  1. Knowl 1 — RMM combines session-level memory construction with online retrieval refinement

    model/method

    Reflective Memory Management (RMM) is a personalized dialogue-agent framework with a memory bank, a base retriever, a learnable reranker, and a response-generating LLM. For each user query, the retriever fetches candidate memories from the bank; the reranker selects a smaller set to provide to the LLM along with the current-session dialogue. The LLM generates a response and citations indicating which retrieved memories supported it. RMM uses those citations to update the reranker online, then appends the new query and response to the current session. When the session ends, RMM extracts topic-based memories from it, adds or merges them into the memory bank, and clears the current-session history. The default setup uses Contriever as the base retriever, retrieves K=20K=20 candidates, and supplies M=5M=5 reranked memories to the LLM.

  2. Knowl 2 — Prospective Reflection builds and updates topic-based memories

    model/method

    Prospective Reflection organizes past dialogue into memories based on semantically coherent topics, rather than relying on fixed turn or session boundaries. A topic may span one or more turns and is represented by a concise summary paired with the raw dialogue segment or segments where the topic was discussed. At the end of each session, an LLM extracts topic summaries and their corresponding dialogue snippets. For each extracted memory, RMM retrieves the KK most semantically similar existing memories from the memory bank and asks an LLM whether to add the new memory as a distinct topic or merge it with a related memory into an updated representation. This process is intended to consolidate related or evolving information while retaining the dialogue evidence associated with each topic.

  3. Knowl 3 — The reranker adapts query and memory embeddings before stochastic selection

    model/method

    RMM's reranker refines the base retriever's top-KK candidates without changing the retriever itself. Let qq be the query embedding and mim_i the embedding of candidate memory ii, for i=1,…,Ki=1,\ldots,K; let WqW_q and WmW_m be learned, dimension-compatible linear transformations. The reranker adapts both embeddings with residual connections and scores each candidate by a dot product:

    q′=q+Wqq,mi′=mi+Wmmi,si=q′⊤mi′.q' = q + W_q q, \qquad m_i' = m_i + W_m m_i, \qquad s_i = {q'}^\top m_i'.

    To make selection stochastic and differentiable, it adds independent Gumbel noise to each score. For independent ui∼Uniform⁡(0,1)u_i \sim \operatorname{Uniform}(0,1), define gi=−log⁡(−log⁡ui)g_i=-\log(-\log u_i) and s~i=si+gi\tilde{s}_i=s_i+g_i. The candidate selection probabilities are

    pi=exp⁡(s~i/τ)∑j=1Kexp⁡(s~j/τ),p_i = \frac{\exp(\tilde{s}_i/\tau)}{\sum_{j=1}^{K}\exp(\tilde{s}_j/\tau)},

    where τ>0\tau>0 is the temperature: lower values make selection more concentrated on high-scoring candidates, while higher values increase stochasticity. RMM selects the top-MM memories for response generation. The implementation describes the reranker as an MLP with a residual connection.

  4. Knowl 4 — LLM citations provide rewards for online REINFORCE updates

    model/method

    To obtain retrieval feedback without labeled relevance data, RMM prompts the response-generating LLM to produce both an answer and citations to the retrieved memories that support it in a single call. A retrieved memory cited in the final response receives reward +1+1; a retrieved memory not cited receives reward −1-1. The reranker parameters ϕ\phi are updated with REINFORCE according to

    Δϕ=η(R−b)∇ϕlog⁡P(MM∣q,MK;ϕ),\Delta \phi = \eta (R-b)\nabla_{\phi}\log P(M_M\mid q,M_K;\phi),

    where qq is the query, MKM_K is the base retriever's candidate set, MMM_M is the reranker's selected set, RR is the citation-derived reward, bb is a baseline, and η\eta is the policy-gradient learning rate. The reported training setup uses batch size 4, Gumbel temperature τ=0.5\tau=0.5, baseline b=0.5b=0.5, and learning rate η=1×10−3\eta=1\times10^{-3}. The feedback is based on whether the generator actually uses retrieved evidence, rather than on human relevance labels.

  5. Knowl 5 — RMM improves retrieval and response metrics on MSC and LongMemEval

    data/table

    The evaluation compares memory and retrieval methods on MSC and LongMemEval. MSC response quality is measured with METEOR and BERTScore; LongMemEval uses Recall@5 and answer accuracy. Scores are percentages averaged over three runs. RMM outperforms the corresponding RAG baseline for every listed retriever and metric; with GTE, it gains 5.9 percentage points in MSC METEOR, 5.0 in MSC BERTScore, 7.4 in LongMemEval Recall@5, and 6.8 in LongMemEval accuracy relative to RAG with GTE. The oracle row is available only for LongMemEval.

    Method Retriever MSC METEOR MSC BERT LongMemEval Recall@5 LongMemEval Acc.
    No History – 5.2 10.6 – 0.0
    Long Context – 14.8 31.9 – 57.4
    RAG Contriever 24.8 50.8 54.3 58.8
    RAG Stella 26.2 51.6 59.2 61.4
    RAG GTE 27.5 52.1 62.4 63.6
    MemoryBank Specific 20.1 40.3 58.6 59.6
    LD-Agent Specific 25.4 51.5 56.8 59.2
    RMM Contriever 30.8 55.4 60.4 61.2
    RMM Stella 31.9 56.3 65.9 64.8
    RMM GTE 33.4 57.1 69.8 70.4
    RAG Oracle Oracle retrieval – – 100.0 90.2

    The no-history and long-context results show the value of history and the limits of using a fixed context window alone. RMM's gains over RAG show that memory organization and adaptive reranking add value beyond base retrieval.

  6. Knowl 6 — Ablations show gains from both reflection mechanisms and the reranker

    data/table

    The ablation compares components using Contriever and Gemini-1.5-Flash, with percentage scores on MSC (METEOR and BERTScore) and LongMemEval (Recall@5 and accuracy). Adding Prospective Reflection (PR) to RAG improves all four metrics. Retrospective Reflection (RR) without a reranker performs poorly when the retriever itself is fine-tuned instead. Adding RR with the reranker improves over RAG, and the complete RMM combination performs best on all four reported metrics.

    Variant MSC METEOR MSC BERT LongMemEval Recall@5 LongMemEval Acc.
    RAG 24.8 50.8 54.3 58.8
    RAG + PR 28.6 53.3 57.4 59.6
    RR without reranker 20.3 31.8 34.2 31.0
    RAG + RR with reranker 27.5 52.2 58.8 60.2
    RMM 30.8 55.4 60.4 61.2

    The authors attribute the poor no-reranker result to directly fine-tuning the retriever with RL rewards, which requires more training data and may cause catastrophic forgetting when data is insufficient.

  7. Knowl 7 — Topic-adaptive granularity approaches per-instance oracle performance

    empirical result

    A granularity comparison on 100 randomly sampled LongMemEval instances used GTE retrieval and Gemini-1.5-Flash generation. It measured passage retrieval with Recall@5 and question answering with accuracy. The proposed Prospective Reflection (PR) granularity outperformed fixed turn, fixed session, and mixed turn-and-session retrieval on both metrics. Selecting the better of turn or session granularity separately for each instance (“Best”) performed better still, so PR approaches but does not reach this per-instance oracle.

    Granularity Passage retrieval Recall@5 (%) Question-answering accuracy (%)
    Turn 47 29
    Session 69 34
    Mixed turn and session 38 17
    Prospective Reflection (PR) 78 49
    Best per-instance granularity 86 58

    The mixed pool performs worse than either turn or session retrieval in this comparison; the authors suggest that mixing granularities can introduce noise. PR instead groups fragmented dialogue into coherent memory structures.

  8. Knowl 8 — Citation-based usefulness labels agree with an LLM judge

    empirical result

    On LongMemEval, Gemini-1.5-Pro judged whether memories cited by the generator were useful for producing the response. The judge's labels were compared with RMM's citation-derived useful/not-useful classifications. Overall precision, recall, and F1 were 87.6%, 85.8%, and 86.7%, respectively. The results support citations as an imperfect but informative signal for the reranker's reward assignment.

    Memory class Precision (%) Recall (%) F1 (%)
    Useful memory 89.4 91.1 90.2
    Not useful memory 87.2 84.6 85.9
    Overall 87.6 85.8 86.7
  9. Knowl 9 — Increasing the number of retrieved and reranked memories improves LongMemEval results

    data/table

    The reported LongMemEval comparison tests a smaller configuration (Top-K=20K=20, Top-M=5M=5) against a larger one (Top-K=50K=50, Top-M=10M=10). It reports Recall@5 and accuracy for the smaller setting, and Recall@10 and accuracy for the larger setting. Increasing the candidate and selected-memory counts improves both retrieval recall and answer accuracy for each tested retriever. GTE achieves the strongest results in both configurations.

    Retriever Recall@5 (%) Acc. (%) Recall@10 (%) Acc. (%)
    Contriever 60.4 61.2 67.2 66.8
    Stella 65.9 64.8 70.6 71.0
    GTE 69.8 70.4 74.4 73.8

    For GTE, the larger setting raises accuracy from 70.4% to 73.8% and recall from 69.8% to 74.4%.

  10. Knowl 10 — RMM has computational, modality, and memory-update limitations

    limitation

    The paper identifies three limitations of RMM: reinforcement-learning-based reranking can be computationally expensive, particularly at large scale or in real-time applications; the current framework is primarily text-based and therefore does not directly cover dialogue involving images, audio, or video; and the memory-update mechanism may need further optimization for efficiently handling dynamically evolving long-term interactions. The paper proposes exploring more efficient RL and lightweight reranking, extending the framework to multimodal dialogue, and developing privacy-preserving approaches for deployment.

Coverage note — Omitted the secondary generator-model comparison, offline supervised retriever pretraining, RL convergence plot, and MSC human-judge agreement results because they are supporting analyses rather than additional core mechanisms or benchmark findings.

References

  1. 1.Sanghwan Bae, Donghyun Kwak, Soyoung Kang, Min Young Lee, Sungdong Kim, Yuin Jeong, Hyeri Kim, Sang-Woo Lee, Woomyoung Park, and Nako Sung. 2022. Keep me updated! memory management in long-term conversations. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3769–3787, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Jan Buchmann, Xiao Liu, and Iryna Gurevych. 2024. Attribute or abstain: Large language models as long document assistants. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8113–8140, Miami, Florida, USA. Association for Computational Linguistics.
  4. 4.Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, Kai Zheng, Defu Lian, and Enhong Chen. 2024. When large language models meet personalization: perspectives of challenges and opportunities. World Wide Web, 27(4).
  5. 5.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. 2025. Mem0: Building production-ready ai agents with scalable long-term memory. Preprint, arXiv:2504.19413.
  6. 6.Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can LLM be a personalized judge? In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 10126–10141, Miami, Florida, USA. Association for Computational Linguistics.
  7. 7.Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, and Chenyi Zhuang. 2024. Intelligent agents with llm-based process automation. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, page 5018–5027, New York, NY, USA. Association for Computing Machinery.
  8. 8.E.J. Gumbel. 1954. Statistical Theory of Extreme Values and Some Practical Applications: A Series of Lectures. Applied mathematics series. U.S. Government Printing Office.
  9. 9.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  10. 10.Eric Jang, Shixiang Gu, and Ben Poole. 2017. Categorical reparameterization with gumbel-softmax. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  11. 11.Zhouyu Jiang, Mengshu Sun, Lei Liang, and Zhiqiang Zhang. 2024. Retrieve, summarize, plan: Advancing multi-hop question answering with an iterative approach. ArXiv preprint, abs/2407.13101.
  12. 12.Krishnaram Kenthapadi, Mehrnoosh Sameki, and Ankur Taly. 2024. Grounding and evaluation for large language models: Practical challenges and lessons learned (survey). In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, page 6523–6533, New York, NY, USA. Association for Computing Machinery.
  13. 13.Seo Hyun Kim, Kai Tzu-iunn Ong, Taeyoon Kwon, Namyoung Kim, Keummin Ka, SeongHyeon Bae, Yohan Jo, Seung-won Hwang, Dongha Lee, and Jinyoung Yeo. 2024. Theanine: Revisiting memory management in long-term conversations with timeline-augmented response generation. ArXiv preprint, abs/2406.10996.
  14. 14.Saydulu Kolasani. 2023. Optimizing natural language processing, large language models (llms) for efficient customer service, and hyper-personalization to enable sustainable growth and revenue. Transactions on Latest Trends in Artificial Intelligence, 4(4).
  15. 15.Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, and Kangwook Lee. 2023. Prompted LLMs as chatbot modules for long open-domain conversation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 4536–4554, Toronto, Canada. Association for Computational Linguistics.
  16. 16.Hao Li, Chenghao Yang, An Zhang, Yang Deng, Xiang Wang, and Tat-Seng Chua. 2024a. Hello again! llm-powered personalized agent for long-term dialogue. ArXiv preprint, abs/2406.05925.
  17. 17.Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, and Yixuan Su. 2024b. Alr2: A retrieve-then-reason framework for long-context question answering. ArXiv preprint, abs/2410.03227.
  18. 18.Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al. 2024c. Personal llm agents: Insights and survey about the capability, efficiency and security. ArXiv preprint, abs/2401.05459.
  19. 19.Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, et al. 2024d. Scbench: A kv cache-centric analysis of long-context methods. ArXiv preprint, abs/2412.10319.
  20. 20.Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards general text embeddings with multi-stage contrastive learning. ArXiv preprint, abs/2308.03281.
  21. 21.Hao Liu, Matei Zaharia, and Pieter Abbeel. 2024a. Ringattention with blockwise transformers for near-infinite context. In The Twelfth International Conference on Learning Representations.
  22. 22.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024b. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
  23. 23.Xiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li, Bo Li, Xuming Hu, and Xiaowen Chu. 2025. Chunkkv: Semantic-preserving kv cache compression for efficient long-context llm inference. ArXiv preprint, abs/2502.00299.
  24. 24.Junru Lu, Siyu An, Mingbao Lin, Gabriele Pergola, Yulan He, Di Yin, Xing Sun, and Yunsheng Wu. 2023. Memochat: Tuning llms to use memos for consistent long-range open-domain conversation. ArXiv preprint, abs/2308.08239.
  25. 25.Chris J Maddison, Daniel Tarlow, and Tom Minka. 2014. A* sampling. In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc.
  26. 26.Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024. Evaluating very long-term conversational memory of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13851–13870, Bangkok, Thailand. Association for Computational Linguistics.
  27. 27.Michael McCloskey and Neal J. Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. volume 24 of Psychology of Learning and Motivation, pages 109–165. Academic Press.
  28. 28.John Mendonça, Alon Lavie, and Isabel Trancoso. 2024. On the benchmarking of LLMs for open-domain dialogue evaluation. In Proceedings of the 6th Workshop on NLP for Conversational AI (NLP4ConvAI 2024), pages 1–12, Bangkok, Thailand. Association for Computational Linguistics.
  29. 29.Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2024. Memgpt: Towards llms as operating systems. Preprint, arXiv:2310.08560.
  30. 30.Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo, Hao Cheng, Dongsheng Li, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Jianfeng Gao. 2025. Secom: On memory construction and retrieval for personalized conversational agents. In The Thirteenth International Conference on Learning Representations.
  31. 31.Jiahuan Pei, Pengjie Ren, and Maarten de Rijke. 2021. A cooperative memory network for personalized task-oriented dialogue systems with incomplete user profiles. In WWW ’21: The Web Conference 2021, Virtual Event / Ljubljana, Slovenia, April 19-23, 2021, pages 1552–1561. ACM / IW3C2.
  32. 32.Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. 2025. Zep: A temporal knowledge graph architecture for agent memory. Preprint, arXiv:2501.13956.
  33. 33.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 31210–31227. PMLR.
  34. 34.Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Wei-Lin Chen, Chao-Wei Huang, Yu Meng, and Yun-Nung Chen. 2024. Two tales of persona in LLMs: A survey of role-playing and personalization. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 16612–16631, Miami, Florida, USA. Association for Computational Linguistics.
  35. 35.Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997. PARADISE: A framework for evaluating spoken dialogue agents. In 35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Computational Linguistics, pages 271–280, Madrid, Spain. Association for Computational Linguistics.
  36. 36.Qingyue Wang, Liang Ding, Yanan Cao, Zhiliang Tian, Shi Wang, Dacheng Tao, and Li Guo. 2023. Recursively summarizing enables long-term dialogue memory in large language models. ArXiv preprint, abs/2308.15022.
  37. 37.Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. 2024. Agent workflow memory. Preprint, arXiv:2409.07429.
  38. 38.Joseph Weizenbaum. 1966. Eliza—a computer program for the study of natural language communication between man and machine. Commun. ACM, 9(1):36–45.
  39. 39.Qingsong Wen, Jing Liang, Carles Sierra, Rose Luckin, Richard Tong, Zitao Liu, Peng Cui, and Jiliang Tang. 2024. Ai for education (ai4edu): Advancing personalized education with llm and adaptive learning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, page 6743–6744, New York, NY, USA. Association for Computing Machinery.
  40. 40.S. Whittaker, Q. Jones, and L. Terveen. 2002. Managing long term communications: conversation and contact management. In Proceedings of the 35th Annual Hawaii International Conference on System Sciences, pages 1070–1079.
  41. 41.Michael David Williams and James D. Hollan. 1981. The process of retrieval from very long-term memory. Cognitive Science, 5(2):87–119.
  42. 42.Ronald J. Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn., 8(3–4):229–256.
  43. 43.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. 2024. Longmemeval: Benchmarking chat assistants on long-term interactive memory. ArXiv preprint, abs/2410.10813.
  44. 44.Jing Xu, Arthur Szlam, and Jason Weston. 2022. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197, Dublin, Ireland. Association for Computational Linguistics.
  45. 45.Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, and Yongfeng Zhang. 2025. A-mem: Agentic memory for llm agents. Preprint, arXiv:2502.12110.
  46. 46.Chiyu Zhang, Yifei Sun, Jun Chen, Jie Lei, Muhammad Abdul-Mageed, Sinong Wang, Rong Jin, Sem Park, Ning Yao, and Bo Long. 2024a. Spar: Personalized content-based recommendation via long engagement attention. ArXiv preprint, abs/2402.10555.
  47. 47.Dun Zhang, Jiacheng Li, Ziyang Zeng, and Ful ong Wang. 2024b. Jasper and stella: distillation of sota embedding models. ArXiv preprint, abs/2412.19048.
  48. 48.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  49. 49.Zhehao Zhang, Ryan A Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, et al. 2024c. Personalization of large language models: A survey. ArXiv preprint, abs/2411.00027.
  50. 50.Zheyuan Zhang, Daniel Zhang-Li, Jifan Yu, Linlu Gong, Jinchang Zhou, Zhiyuan Liu, Lei Hou, and Juanzi Li. 2024d. Simulating classroom education with llm-empowered agents. ArXiv preprint, abs/2406.19226.
  51. 51.Liang Zhao, Xiachong Feng, Xiaocheng Feng, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin, and Ting Liu. 2024. Length extrapolation of transformers: A survey from the perspective of positional encoding. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9959–9977, Miami, Florida, USA. Association for Computational Linguistics.
  52. 52.Chuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang, Jingyao Li, Jing Xiong, Xiaozhe Ren, Michael Ng, Xin Jiang, Zhenguo Li, and Yu Li. 2024. Dape: Data-adaptive positional encoding for length extrapolation. In Advances in Neural Information Processing Systems, volume 37, pages 26659–26700. Curran Associates, Inc.
  53. 53.Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2024. Memorybank: Enhancing large language models with long-term memory. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada, pages 19724–19731. AAAI Press.

Citation

MLA
Tan, Z., et al. “In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 8416–39, https://doi.org/10.18653/v1/2025.acl-long.413.
APA
Tan, Z., Yan, J., Hsu, I.-H., Han, R., Wang, Z., Le, L., Song, Y., Chen, Y., Palangi, H., Lee, G., Iyer, A. R., Chen, T., Liu, H., Lee, C.-Y., & Pfister, T. (2025). In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8416–8439. https://doi.org/10.18653/v1/2025.acl-long.413
Chicago
Tan, Z., J. Yan, I.-H. Hsu, et al. 2025. “In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8416–39. https://doi.org/10.18653/v1/2025.acl-long.413.
Harvard
Tan, Z. et al. (2025) “In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8416–8439. Available at: https://doi.org/10.18653/v1/2025.acl-long.413.
Vancouver
1. Tan Z, Yan J, Hsu I-H, et al (2025) In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8416–8439

BibTeX

@inproceedings{tan-etal-2025-prospect,
    title = "In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents",
    author = "Tan, Zhen  and
      Yan, Jun  and
      Hsu, I-Hung  and
      Han, Rujun  and
      Wang, Zifeng  and
      Le, Long  and
      Song, Yiwen  and
      Chen, Yanfei  and
      Palangi, Hamid  and
      Lee, George  and
      Iyer, Anand Rajan  and
      Chen, Tianlong  and
      Liu, Huan  and
      Lee, Chen-Yu  and
      Pfister, Tomas",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.413/",
    doi = "10.18653/v1/2025.acl-long.413",
    pages = "8416--8439",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/