Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
Hao LiChenghao YangAn ZhangYang DengXiang WangTat-Seng Chua
Proposes LD-Agent, a model-agnostic framework that integrates modular event memory banks and dynamic user-agent persona extraction to maintain consistency and context across multi-session conversations.
Modern conversational artificial intelligence systems frequently struggle to sustain long-term engagement across multiple interaction sessions. While large language models excel at brief, single-session exchanges, practical applications—such as personalized companions, virtual assistants, and customer support—demand that agents remember past events, track conversational history over days or years, and preserve consistent personal traits. Most existing systems either isolate event memory from user profiling or rely heavily on rigid architectures that fail to generalize across diverse operational domains.
The article introduces and evaluates the Long-term Dialogue Agent (LD-Agent), a modular, model-agnostic framework designed to deliver coherent, personalized interactions across extended multi-session dialogues. The primary objective is to demonstrate that coordinating dynamic event perception with bidirectional persona modeling significantly enhances conversational quality, character consistency, and domain adaptability.
To achieve this, the authors designed a three-part modular system comprising an event memory perception unit, a persona extraction unit, and a response generation unit. The event memory perception unit separates historical session summaries in a long-term memory bank from ongoing dialogue context in a short-term memory cache, employing a retrieval mechanism based on semantic relevance, topic noun overlap, and time decay. The persona extraction unit dynamically extracts and updates personality traits for both the user and the agent using fine-tuned instruction models. The authors evaluated the framework on standard five-session benchmarks (MSC and Conversation Chronicles), tested cross-domain transfers between human-annotated and synthetic datasets, assessed multiparty dialogue on the Ubuntu IRC benchmark, and conducted human evaluations with eight evaluators.
The evaluation produced four primary findings. First, integrating the proposed framework led to substantial performance improvements across all tested baseline models; for example, on the Conversation Chronicles dataset, fine-tuning an open-source model with the framework increased generation overlap metrics by approximately 50% to 60% compared to standard tuning. Second, the framework outperformed the previous state-of-the-art architecture across multi-session benchmarks. Third, human assessments indicated that topic-based memory retrieval substantially exceeded direct semantic retrieval in both accuracy and recall, translating into marked gains in conversation coherence, fluency, and user engagement. Fourth, cross-domain evaluations revealed strong generalization, as models trained on one dataset retained high performance when evaluated on another with distinct collection properties.
These findings indicate that dialogue systems can achieve long-term coherence without requiring monolithic, bespoke architectures. Organizations deploying conversational systems can adapt existing language models across tasks—such as direct one-on-one chats and complex multiparty conversations—while keeping memory storage and persona management computationally efficient through modular tuning.
Organizations developing customer-facing or companion systems should consider adopting modular memory and persona tracking architectures to improve retention and response quality. Technical teams should implement topic-aware and recency-weighted retrieval rather than relying solely on raw semantic search. Prior to large-scale deployment, practitioners should conduct pilot tests in realistic environments to validate conversational dynamics.
The primary limitation noted in the article is the reliance on synthetic or crowd-sourced multi-session datasets, which may not capture all complexities of authentic real-world dialogues. Additionally, the individual modules currently employ basic summarization and extraction mechanisms, leaving room for further optimization. Consequently, while the framework demonstrates strong technical validity, stakeholders should expect further performance variations when migrating to live production environments.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Its LOCOMO benchmark formalizes long-term conversational memory and event reasoning, providing a direct evaluation foundation for LD-Agent’s memory design and claims.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This study establishes explicit persona conditioning as a way to improve dialogue consistency, a prerequisite for understanding LD-Agent’s dynamic user and agent persona modeling.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Its persistent experience memory and reflection architecture introduce core ideas behind agents that use accumulated history to guide coherent future behavior.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Its taxonomy of agent profiling and hybrid short- and long-term memory clarifies the architectural concepts that LD-Agent specializes for personalized dialogue.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). It advances long-history dialogue memory beyond retrieval by iteratively reconstructing relevant evidence through guided graph traversal.
- Paper: Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory, Xin Yu et al. (2026). It extends long-term conversational memory with hierarchical graph exploration for queries that require broad, multi-hop evidence across conversation histories.
- Paper: Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models, Harshita Chopra et al. (2026). It continues personalized memory retrieval by using simulated future actions to uncover relevant history that direct similarity search misses.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). It develops persistent agent memory further with consolidation, adaptive forgetting, and reconsolidation across extended interaction streams.
