Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations
Jihyoung JangMinseong BooHyounghun Kim
Introduces Conversation Chronicles, a 1-million multi-session dialogue dataset incorporating varied time intervals and speaker relationships, alongside ReBot, an efficient 630M-parameter model that maintains long-term conversational context across consecutive interactions.
Most modern conversational artificial intelligence systems are designed for short, single-session interactions. Because they fail to retain long-term context across multiple conversations occurring over extended time periods, they cannot maintain coherent, ongoing relationships with users. The article addresses this challenge by establishing a framework for long-term dialogue that explicitly incorporates elapsed time between sessions and defined social relationships between speakers.
The main objective of the article is to create a large-scale, multi-session dialogue benchmark and demonstrate an efficient, compact dialogue model capable of sustaining context, recalling past events, and adapting conversational dynamics based on explicit speaker relationships and temporal gaps.
To accomplish this, the authors built an event graph using natural language inference to connect related narrative events and queried a large language model with structured prompts containing ten distinct human relationships and five time intervals ranging from hours to years. This yielded CONVERSATION CHRONICLES, a dataset consisting of one million multi-session dialogues across 200,000 five-session episodes. Using this dataset, the authors developed REBOT, a conversational model with approximately 630 million parameters comprising an event summarization module based on T5 and a dialogue generation module based on BART. Rigorous human evaluations—including professional third-party assessments, consensus scoring, and live interactive chat trials—were conducted to evaluate the dataset and model.
Human evaluations established that CONVERSATION CHRONICLES achieved high quality, scoring 4.33 out of 5 overall and consistently outperforming the benchmark multi-session dataset MSC across coherence, consistency, and temporal logic. Model assessments revealed that REBOT achieved strong generation performance, scoring 4.78 for engagingness, 4.74 for humanness, and 4.14 for memorability. In live human interactions, REBOT outperformed the much larger 2.7-billion-parameter MSC model across all criteria, earning an overall score of 4.23 compared to MSC's 2.93 (representing an improvement of roughly 44%). Ablation experiments confirmed that removing temporal or relational context resulted in generic time references and inconsistent speaker personas.
These findings show that conversational AI performance over long time horizons does not strictly require massive parameter scales; instead, parameter-efficient architectures can achieve superior contextual retention and user engagement when trained on structured multi-session data. Incorporating explicit relationships and temporal intervals allows conversational agents to adapt their tone, recall relevant history, and navigate interpersonal shifts over time, which improves user trust and interaction quality.
Organizations deploying automated dialogue systems—such as virtual counseling, tutoring, or ongoing customer support—should adopt structured chronological summarization and explicit relational framing. While the findings strongly support this architectural approach, development teams should proceed with caution regarding the current boundary conditions, as the study evaluated a fixed set of ten relationships, preset time categories, and data generated from a single large language model family. Future initiatives should pilot the framework across broader relationship domains, test diverse base model architectures, and implement strict safety controls against inappropriate domain-specific advice.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). Its persona-conditioned dialogue experiments establish how persistent speaker profiles can improve conversational consistency, a useful foundation for the source’s explicit modeling of social relationships.
- Paper: Hello Again! LLM-powered Personalized Agent for Long-term Dialogue, Hao Li et al. (2025). Building on the source’s multi-session benchmark, this work evaluates a modular agent that combines event memory with evolving user and agent personas for more personalized long-term dialogue.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It extends the source’s short multi-session setting into conversations spanning months, testing long-term recall, temporal reasoning, and multimodal memory at greater scale.
- Paper: In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents, Zhen Tan et al. (2025). It carries the source’s long-term dialogue memory problem forward with adaptive topic organization and retrieval that improve recall across multi-session conversations.
- Paper: GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations, Jingbo Yang et al. (2026). It extends relationship-aware memory beyond dyadic dialogue by evaluating whether agents can track speakers, roles, and changing beliefs across group conversations.
