Hello Again! LLM-powered Personalized Agent for Long-term Dialogue

Hao LiChenghao YangAn ZhangYang DengXiang WangTat-Seng Chua

article2025NAACL162 citations

Proposes LD-Agent, a model-agnostic framework that integrates modular event memory banks and dynamic user-agent persona extraction to maintain consistency and context across multi-session conversations.

Listen

Modern conversational artificial intelligence systems frequently struggle to sustain long-term engagement across multiple interaction sessions. While large language models excel at brief, single-session exchanges, practical applications—such as personalized companions, virtual assistants, and customer support—demand that agents remember past events, track conversational history over days or years, and preserve consistent personal traits. Most existing systems either isolate event memory from user profiling or rely heavily on rigid architectures that fail to generalize across diverse operational domains.

The article introduces and evaluates the Long-term Dialogue Agent (LD-Agent), a modular, model-agnostic framework designed to deliver coherent, personalized interactions across extended multi-session dialogues. The primary objective is to demonstrate that coordinating dynamic event perception with bidirectional persona modeling significantly enhances conversational quality, character consistency, and domain adaptability.

To achieve this, the authors designed a three-part modular system comprising an event memory perception unit, a persona extraction unit, and a response generation unit. The event memory perception unit separates historical session summaries in a long-term memory bank from ongoing dialogue context in a short-term memory cache, employing a retrieval mechanism based on semantic relevance, topic noun overlap, and time decay. The persona extraction unit dynamically extracts and updates personality traits for both the user and the agent using fine-tuned instruction models. The authors evaluated the framework on standard five-session benchmarks (MSC and Conversation Chronicles), tested cross-domain transfers between human-annotated and synthetic datasets, assessed multiparty dialogue on the Ubuntu IRC benchmark, and conducted human evaluations with eight evaluators.

The evaluation produced four primary findings. First, integrating the proposed framework led to substantial performance improvements across all tested baseline models; for example, on the Conversation Chronicles dataset, fine-tuning an open-source model with the framework increased generation overlap metrics by approximately 50% to 60% compared to standard tuning. Second, the framework outperformed the previous state-of-the-art architecture across multi-session benchmarks. Third, human assessments indicated that topic-based memory retrieval substantially exceeded direct semantic retrieval in both accuracy and recall, translating into marked gains in conversation coherence, fluency, and user engagement. Fourth, cross-domain evaluations revealed strong generalization, as models trained on one dataset retained high performance when evaluated on another with distinct collection properties.

These findings indicate that dialogue systems can achieve long-term coherence without requiring monolithic, bespoke architectures. Organizations deploying conversational systems can adapt existing language models across tasks—such as direct one-on-one chats and complex multiparty conversations—while keeping memory storage and persona management computationally efficient through modular tuning.

Organizations developing customer-facing or companion systems should consider adopting modular memory and persona tracking architectures to improve retention and response quality. Technical teams should implement topic-aware and recency-weighted retrieval rather than relying solely on raw semantic search. Prior to large-scale deployment, practitioners should conduct pilot tests in realistic environments to validate conversational dynamics.

The primary limitation noted in the article is the reliance on synthetic or crowd-sourced multi-session datasets, which may not capture all complexities of authentic real-world dialogues. Additionally, the individual modules currently employ basic summarization and extraction mechanisms, leaving room for further optimization. Consequently, while the framework demonstrates strong technical validity, stakeholders should expect further performance variations when migrating to live production environments.

Cover for Hello Again! LLM-powered Personalized Agent for Long-term Dialogue

Abstract

Open-domain dialogue systems have seen remarkable advancements with the development of large language models (LLMs). Nonetheless, most existing dialogue systems predominantly focus on brief single-session interactions, neglecting the real-world demands for long-term companionship and personalized interactions with chatbots. Crucial to addressing this real-world need are event summary and persona management, which enable reasoning for appropriate long-term dialogue responses. Recent progress in the human-like cognitive and reasoning capabilities of LLMs suggests that LLM-based agents could significantly enhance automated perception, decision-making, and problem-solving. In response to this potential, we introduce a model-agnostic framework, the Long-term Dialogue Agent (LD-Agent), which incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation. For the event memory module, long and short-term memory banks are employed to separately focus on historical and ongoing sessions, while a topic-based retrieval mechanism is introduced to enhance the accuracy of memory retrieval. Furthermore, the persona module conducts dynamic persona modeling for both users and agents. The integration of retrieved memories and extracted personas is subsequently fed into the generator to induce appropriate responses. The effectiveness, generality, and cross-domain capabilities of LD-Agent are empirically demonstrated across various illustrative benchmarks, models, and tasks. The code is released at https://github.com/leolee99/LD-Agent.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Task Definition
  • 3.2 Event Perception
  • 3.2.1 Long-term Memory
  • 3.2.2 Short-term Memory
  • 3.3 Dynamic Personas Extraction
  • 3.4 Response Generation
  • 4 Experiments
  • 4.1 Evaluation Settings
  • 4.2 Evaluation Pipeline
  • 4.3 Results of Multi-Session Dialogue
  • 4.4 Ablation Studies
  • 4.5 Persona Extraction Analysis
  • 4.6 Human Evaluation
  • 4.7 Generality Analysis
  • 5 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • Appendix
  • A Detailed Evaluation Settings
  • A.1 Datasets
  • A.2 Metrics
  • A.3 Baselines
  • A.4 Implementation Details
  • B Qualitative Analysis
  • B.1 Persona Ablation
  • B.2 Memory Ablation
  • B.3 Event Summarizer Analysis
  • B.4 Persona Extractor Analysis
  • B.5 Response Generation Analysis
  • C Additional Experimental Results
  • C.1 Part of Speech Importance Analysis
  • C.2 Generation Diversity Analysis
  • D Prompt
  • D.1 Prompt of Event Summary
  • D.2 Prompt of Persona Extraction
  • D.3 Prompt of Response Generation

Knowls

  1. Knowl 1 — Paper content unavailable for knowl extraction

    limitation

    No knowls from the paper's own contribution could be produced because the document content was unavailable; any substantive claims, methods, or results would require access to the paper's actual text.

Coverage note — The content of the attached paper was not accessible in the provided input, so no contributed material (methods, theory, experiments, or results) could be extracted or deliberately omitted; the single entry below records this extraction failure.

References

  1. 1.[1] A. Krizhevsky, I. Sutskever, and G. E. Hinton, "ImageNet classification with deep convolutional neural networks," in Advances in Neural Information Processing Systems (NIPS), 2012, pp. 1097–1105.
  2. 2.[2] K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," in International Conference on Learning Representations (ICLR), 2015.
  3. 3.[3] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.
  4. 4.[4] S. Zagoruyko and N. Komodakis, "Wide residual networks," in Proceedings of the British Machine Vision Conference (BMVC), 2016.
  5. 5.[5] G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, "Densely connected convolutional networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4700–4708.
  6. 6.[6] J. Hu, L. Shen, and G. Sun, "Squeeze-and-excitation networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7132–7141.
  7. 7.[7] S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, "Aggregated residual transformations for deep neural networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1492–1500.
  8. 8.[8] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, "MobileNets: Efficient convolutional neural networks for mobile vision applications," arXiv preprint arXiv:1704.04861, 2017.
  9. 9.[9] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "MobileNetV2: Inverted residuals and linear bottlenecks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510–4520.
  10. 10.[10] F. Chollet, "Xception: Deep learning with depthwise separable convolutions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1251–1258.
  11. 11.[11] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 1–9.
  12. 12.[12] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi, "Inception-v4, inception-resnet and the impact of residual connections on learning," in Proceedings of the AAAI Conference on Artificial Intelligence, 2017.
  13. 13.[13] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, "An image is worth 16x16 words: Transformers for image recognition at scale," in International Conference on Learning Representations (ICLR), 2021.
  14. 14.[14] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, "A convnet for the 2020s," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 11976–11986.
  15. 15.[15] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016.
  16. 16.[16] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," in International Conference on Learning Representations (ICLR), 2015.
  17. 17.[17] I. Loshchilov and F. Hutter, "SGDR: Stochastic gradient descent with warm restarts," in International Conference on Learning Representations (ICLR), 2017.
  18. 18.[18] I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," in International Conference on Learning Representations (ICLR), 2019.
  19. 19.[19] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, "PyTorch: An imperative style, high-performance deep learning library," in Advances in Neural Information Processing Systems (NIPS), 2019, pp. 8024–8035.
  20. 20.[20] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, "ImageNet large scale visual recognition challenge," International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015.
  21. 21.[21] A. Coates, A. Y. Ng, and H. Lee, "An analysis of single-layer networks in unsupervised feature learning," in Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2011, pp. 215–223.
  22. 22.[22] A. Krizhevsky, "Learning multiple layers of features from tiny images," Technical Report, University of Toronto, 2009.
  23. 23.[23] H. Xiao, K. Rasul, and R. Vollgraf, "Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms," arXiv preprint arXiv:1708.07747, 2017.
  24. 24.[24] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, "Gradient-based learning applied to document recognition," Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  25. 25.[25] K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, "Return of the devil in the detail: Delving deep into convolutional nets," in Proceedings of the British Machine Vision Conference (BMVC), 2014.
  26. 26.[26] M. Tan and Q. V. Le, "EfficientNet: Rethinking model scaling for convolutional neural networks," in Proceedings of the 36th International Conference on Machine Learning (ICML), 2019, pp. 6105–6114.
  27. 27.[27] M. Tan, R. Pang, and Q. V. Le, "EfficientDet: Scalable and efficient object detection," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 10781–10790.
  28. 28.[28] B. Zoph and Q. V. Le, "Neural architecture search with reinforcement learning," in International Conference on Learning Representations (ICLR), 2017.
  29. 29.[29] H. Cai, L. Zhu, and S. Han, "ProxylessNAS: Direct neural architecture search on target task and hardware," in International Conference on Learning Representations (ICLR), 2019.
  30. 30.[30] M. Welling and Y. W. Teh, "Bayesian learning via stochastic gradient Langevin dynamics," in Proceedings of the 28th International Conference on Machine Learning (ICML), 2011, pp. 681–688.

Citation

MLA
Li, H., et al. “Hello Again! LLM-powered Personalized Agent for Long-term Dialogue”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 5259–76, https://doi.org/10.18653/v1/2025.naacl-long.272.
APA
Li, H., Yang, C., Zhang, A., Deng, Y., Wang, X., & Chua, T.-S. (2025). Hello Again! LLM-powered Personalized Agent for Long-term Dialogue. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5259–5276. https://doi.org/10.18653/v1/2025.naacl-long.272
Chicago
Li, H., C. Yang, A. Zhang, Y. Deng, X. Wang, and T.-S. Chua. 2025. “Hello Again! LLM-powered Personalized Agent for Long-term Dialogue”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5259–76. https://doi.org/10.18653/v1/2025.naacl-long.272.
Harvard
Li, H. et al. (2025) “Hello Again! LLM-powered Personalized Agent for Long-term Dialogue”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5259–5276. Available at: https://doi.org/10.18653/v1/2025.naacl-long.272.
Vancouver
1. Li H, Yang C, Zhang A, Deng Y, Wang X, Chua T-S (2025) Hello Again! LLM-powered Personalized Agent for Long-term Dialogue. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5259–5276

BibTeX

@inproceedings{li-etal-2025-hello,
    title = "Hello Again! {LLM}-powered Personalized Agent for Long-term Dialogue",
    author = "Li, Hao  and
      Yang, Chenghao  and
      Zhang, An  and
      Deng, Yang  and
      Wang, Xiang  and
      Chua, Tat-Seng",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.272/",
    doi = "10.18653/v1/2025.naacl-long.272",
    pages = "5259--5276",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/