Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations

Jihyoung JangMinseong BooHyounghun Kim

article2023EMNLP63 citations

Introduces Conversation Chronicles, a 1-million multi-session dialogue dataset incorporating varied time intervals and speaker relationships, alongside ReBot, an efficient 630M-parameter model that maintains long-term conversational context across consecutive interactions.

Listen

Most modern conversational artificial intelligence systems are designed for short, single-session interactions. Because they fail to retain long-term context across multiple conversations occurring over extended time periods, they cannot maintain coherent, ongoing relationships with users. The article addresses this challenge by establishing a framework for long-term dialogue that explicitly incorporates elapsed time between sessions and defined social relationships between speakers.

The main objective of the article is to create a large-scale, multi-session dialogue benchmark and demonstrate an efficient, compact dialogue model capable of sustaining context, recalling past events, and adapting conversational dynamics based on explicit speaker relationships and temporal gaps.

To accomplish this, the authors built an event graph using natural language inference to connect related narrative events and queried a large language model with structured prompts containing ten distinct human relationships and five time intervals ranging from hours to years. This yielded CONVERSATION CHRONICLES, a dataset consisting of one million multi-session dialogues across 200,000 five-session episodes. Using this dataset, the authors developed REBOT, a conversational model with approximately 630 million parameters comprising an event summarization module based on T5 and a dialogue generation module based on BART. Rigorous human evaluations—including professional third-party assessments, consensus scoring, and live interactive chat trials—were conducted to evaluate the dataset and model.

Human evaluations established that CONVERSATION CHRONICLES achieved high quality, scoring 4.33 out of 5 overall and consistently outperforming the benchmark multi-session dataset MSC across coherence, consistency, and temporal logic. Model assessments revealed that REBOT achieved strong generation performance, scoring 4.78 for engagingness, 4.74 for humanness, and 4.14 for memorability. In live human interactions, REBOT outperformed the much larger 2.7-billion-parameter MSC model across all criteria, earning an overall score of 4.23 compared to MSC's 2.93 (representing an improvement of roughly 44%). Ablation experiments confirmed that removing temporal or relational context resulted in generic time references and inconsistent speaker personas.

These findings show that conversational AI performance over long time horizons does not strictly require massive parameter scales; instead, parameter-efficient architectures can achieve superior contextual retention and user engagement when trained on structured multi-session data. Incorporating explicit relationships and temporal intervals allows conversational agents to adapt their tone, recall relevant history, and navigate interpersonal shifts over time, which improves user trust and interaction quality.

Organizations deploying automated dialogue systems—such as virtual counseling, tutoring, or ongoing customer support—should adopt structured chronological summarization and explicit relational framing. While the findings strongly support this architectural approach, development teams should proceed with caution regarding the current boundary conditions, as the study evaluated a fixed set of ten relationships, preset time categories, and data generated from a single large language model family. Future initiatives should pilot the framework across broader relationship domains, test diverse base model architectures, and implement strict safety controls against inappropriate domain-specific advice.

Cover for Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations

Abstract

In the field of natural language processing, open-domain chatbots have emerged as an important research topic. However, a major limitation of existing open-domain chatbot research is its singular focus on short single-session dialogue, neglecting the potential need for understanding contextual information in multiple consecutive sessions that precede an ongoing dialogue. Among the elements that compose the context in multi-session conversation settings, the time intervals between sessions and the relationships between speakers would be particularly important. Despite their importance, current research efforts have not sufficiently addressed these dialogical components. In this paper, we introduce a new 1M multi-session dialogue dataset, called CONVERSATION CHRONICLES, for implementing a long-term conversation setup in which time intervals and fine-grained speaker relationships are incorporated. Following recent works, we exploit a large language model to produce the data. The extensive human evaluation shows that dialogue episodes in CONVERSATION CHRONICLES reflect those properties while maintaining coherent and consistent interactions across all the sessions. We also propose a dialogue model, called REBOT, which consists of chronological summarization and dialogue generation modules using only around 630M parameters. When trained on CONVERSATION CHRONICLES, REBOT demonstrates long-term context understanding with a high human engagement score.^1

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Conversation Chronicles
  • 3.1 Event Collection
  • 3.2 Chronological Dynamics
  • 3.3 Conversation Episode Generation
  • 3.4 Quality
  • 4 ReBot
  • 4.1 Chronological Summary
  • 4.2 Dialogue Generation
  • 5 Experiments
  • 5.1 Implementation and Training Details
  • 5.2 Human Evaluation
  • 5.3 Dataset Quality Evaluation
  • 5.4 Comparison to Other Datasets
  • 5.5 Model Performance Evaluation
  • 5.6 Comparison to Other Models
  • 6 Results
  • 7 Conclusion
  • References
  • A Prompts Details
  • B Dataset Filtering Process
  • C Implementation and Training Details
  • D Dialogue Example for Each Relationship
  • E Chronological Summary Example
  • F Ablation Study Example
  • G Generation Quality per Relationship
  • H Episode Example
  • I Human Evaluation

Knowls

  1. Knowl 1 — CONVERSATION CHRONICLES is a million-session, five-session-per-episode dataset

    data/table

    CONVERSATION CHRONICLES is an English multi-session conversation dataset containing 200,000 episodes, each with five sessions, for 1,000,000 sessions total. It contains 11.7 million turns, averaging 11.7 turns per session and 18.03 words per turn. The authors report that this is a larger multi-session dataset than the comparison datasets they tabulate: MSC has 161,000 turns in its training split with up to three sessions and 53,000 turns in the split with up to four sessions; Korean CareCallmem has 160,000 turns across 7,700 episodes. These comparisons establish the scale and turn-length profile of CONVERSATION CHRONICLES.

  2. Knowl 2 — Entailment-linked event graph supplies coherent five-event episode sequences

    model/method

    To create event sequences for multi-session episodes, the authors use narratives from SODA as event descriptions and classify pairs of events with a BERT-base natural-language-inference model fine-tuned on SNLI. They retain event pairs classified as entailment and represent events as nodes in a directed graph, using premise-to-hypothesis ordering on edges to preserve event order and avoid temporal contradictions. They extract possible length-five event sequences from the graph, then remove sequences that share more than three events with another sequence, retaining only one of each such overlapping set.

  3. Knowl 3 — Episodes span five time intervals and ten assigned speaker relationships

    data/table

    Each episode uses one of ten relationships and assigns a time interval between each consecutive pair of sessions. The interval is randomly selected from “a few hours,” “a few days,” “a few weeks,” “a few months,” and “a couple of years.” The reported counts across session-to-session intervals are, respectively, 159,975; 159,928; 160,670; 160,050; and 159,377. The relationship is selected by ChatGPT from the episode’s events rather than assigned randomly. Reported episode counts and shares are: Classmates, 66,090 (33.05%); Neighbors, 49,521 (24.76%); Co-workers, 28,856 (14.43%); Mentee and Mentor, 16,035 (8.02%); Husband and Wife, 13,486 (6.74%); Patient and Doctor, 6,980 (3.49%); Parent and Child, 6,514 (3.26%); Student and Teacher, 5,018 (2.51%); Employee and Boss, 4,811 (2.41%); and Athlete and Coach, 2,689 (1.34%). The authors use approximate interval categories rather than exact durations because they judge minute differences in time units to have little contextual effect.

  4. Knowl 4 — ChatGPT generates sessions conditioned on events, intervals, relationships, and episode history

    model/method

    The authors prompt ChatGPT to generate each session from its event description, the speaker relationship, and the interval since the preceding session, while supplying the events and intervals from earlier sessions as context. The resulting dataset contains 200,000 five-session episodes. They automatically discard an entire episode if any session has more than two speakers, unclear speaker-to-utterance alignment, speakers inconsistent with the assigned relationship, or unnecessary material such as stage directions. They also use a moderation system to remove harmful data. The generation experiments use the ChatGPT model version gpt-3.5-turbo-0301.

  5. Knowl 5 — REBOT compresses session history and conditions response generation on long-term context

    model/method

    REBOT has a chronological summarization module and a dialogue-generation module. The summarizer maps a past session’s dialogue to a concise chronological summary; the generator uses summaries of past sessions together with the current relationship, time interval, and ongoing-session dialogue history. Its next-utterance distribution is P(c∣r,t,s,h)P(c \mid r,t,s,h), where cc is the next utterance, rr is the speaker relationship, tt is the interval since the previous session, ss is a past-session summary, and hh is the utterance history in the current session. The summarizer is T5-base (222 million parameters), and the generator is BART-large (406 million parameters), for approximately 630 million parameters combined. For summarizer training, ChatGPT produced summaries for 100,000 sessions from 20,000 episodes; 80,000 summaries were used for training and 20,000 for validation and testing. REBOT’s dialogue data were split by episode into 160,000 training, 20,000 validation, and 20,000 test episodes. Both modules use AdamW and cross-entropy training. The summarizer used a linear learning-rate scheduler, batch size 32, maximum input length 512 tokens, maximum output length 128 tokens, and at most five epochs with early stopping; training took about six hours on eight NVIDIA RTX A6000 GPUs. The generator used a linear scheduler, batch size 16, maximum input length 1,024 tokens, maximum output length 128 tokens, and at most three epochs with early stopping; training took about three days on the same GPU configuration.

  6. Knowl 6 — Human ratings favor CONVERSATION CHRONICLES over MSC on shared quality criteria

    empirical result

    Evaluators rated 5,000 CONVERSATION CHRONICLES episodes on a 1–5 scale for consistency, coherence, fit to the designated time interval, and fit to the designated relationship. The reported means (standard deviations) were: consistency 4.41 (0.80), coherence 4.04 (1.06), time interval 4.46 (0.77), and relationship 4.40 (1.08); the reported overall mean was 4.33. In a separate comparison, 500 episodes from CONVERSATION CHRONICLES and 500 five-session episodes from MSC were rated on shared criteria, excluding relationship because MSC does not encode one. The reported means for MSC versus CONVERSATION CHRONICLES were 3.87 versus 4.71 for consistency, 3.66 versus 4.51 for coherence, 4.05 versus 4.82 for time-interval fit, and 3.86 versus 4.68 overall. For this comparison, three evaluators rated each episode and their scores were averaged.

  7. Knowl 7 — REBOT-generated episodes receive high human ratings, including for memory

    empirical result

    In a human evaluation of 1,000 REBOT-generated episodes, evaluators rated engagingness at 4.78 (standard deviation 0.54), humanness at 4.74 (0.63), and memorability at 4.14 (0.79), with an overall score of 4.55 on a 1–5 scale. The episodes were generated autoregressively from sampled first sessions, with 100 starting sessions from each of the ten relationship categories; subsequent intervals were randomly chosen. Memorability was intended to assess whether the model correctly retained and recalled information across sessions. Separately, evaluators judged 3,000 summaries—1,000 each from sessions two, three, and four—for whether they fully described the dialogue interaction. The mean summary score was 4.3 out of 5.

  8. Knowl 8 — Interactive evaluation rates REBOT above MSC 2.7B

    empirical result

    In interactive evaluation, in-house evaluators held 50 live chats with each model, with at least six turns per chat. REBOT chats were grounded in an event summary from a previous session; MSC 2.7B chats were grounded in a persona. On a 1–5 scale, the reported means for MSC 2.7B versus REBOT were 3.24 versus 4.06 for engagingness, 2.92 versus 4.52 for humanness, 2.62 versus 4.12 for memorability, and 2.93 versus 4.23 overall. Thus, REBOT received higher average ratings on all four reported measures in this evaluation setup.

  9. Knowl 9 — Ablations indicate that time and relationship inputs affect response specificity and consistency

    empirical result

    The authors report qualitative ablations of REBOT’s temporal and relational inputs. Without time-interval information during training, the model tended to produce generic references to time rather than responses specific to the designated interval. Without relationship information, its responses could fail to maintain a consistent speaker relationship. The paper illustrates these effects with examples but does not report aggregate quantitative ablation scores, so the evidence supports these observed tendencies rather than a measured effect size.

  10. Knowl 10 — The study’s temporal and relational coverage is limited, and outputs depend on the LLM used

    limitation

    The authors identify the fixed set of five interval categories and ten speaker relationships as a limitation that may restrict generalizability. They also note that changing the LLM used for data generation could yield different dialogue types and results even under the same framework. In addition, moderation does not guarantee that all harmful content is removed, and generated responses—including medical advice—may be unsuitable for real-world use; the authors recommend research use or caution in applications.

Coverage note — Omitted the many illustrative dialogue and summary examples because they demonstrate the methods and findings already captured here rather than adding distinct general results.

References

  1. 1.Sanghwan Bae, Donghyun Kwak, Soyoung Kang, Min Young Lee, Sungdong Kim, Yuin Jeong, Hyeri Kim, Sang-Woo Lee, Woomyoung Park, and Nako Sung. 2022. Keep me updated! memory management in long-term conversations. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3769–3787, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  2. 2.Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. In Proceedings of the international AAAI conference on web and social media, volume 14, pages 830–839.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In International Conference on Learning Representations.
  7. 7.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056.
  8. 8.Google. 2023. Google ai updates: Bard and new ai features in search.
  9. 9.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, et al. 2022a. Soda: Million-scale dialogue distillation with social commonsense contextualization. arXiv preprint arXiv:2212.10465.
  10. 10.Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022b. ProsocialDialog: A prosocial backbone for conversational agents. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4005–4029, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  11. 11.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  12. 12.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986–995, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  13. 13.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  14. 14.Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2122–2132, Austin, Texas. Association for Computational Linguistics.
  15. 15.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. ICLR.
  16. 16.Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023. A holistic approach to undesired content detection in the real world. AAAI.
  17. 17.Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017. ParlAI: A dialog research software platform. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 79–84, Copenhagen, Denmark. Association for Computational Linguistics.
  18. 18.OpenAI. 2022. Introducing ChatGPT. https://openai.com/blog/chatgpt.
  19. 19.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  20. 20.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  21. 21.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  22. 22.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  23. 23.Alan Ritter, Colin Cherry, and William B. Dolan. 2011. Data-driven response generation in social media. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, pages 583–593, Edinburgh, Scotland, UK. Association for Computational Linguistics.
  24. 24.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online. Association for Computational Linguistics.
  25. 25.Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, et al. 2022. Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage. arXiv preprint arXiv:2208.03188.
  26. 26.Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020. Can you put it all together: Evaluating conversational agents’ ability to blend skills. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2021–2030, Online. Association for Computational Linguistics.
  27. 27.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca: A strong, replicable instruction-following model. https://crfm.stanford.edu/2023/03/13/alpaca.html.
  28. 28.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  29. 29.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  30. 30.Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu, and Pascale Fung. 2020. Getting to know you: User attribute extraction from dialogues. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 581–589, Marseille, France. European Language Resources Association.
  31. 31.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. arXiv preprint arXiv:2304.01196.
  32. 32.Jing Xu, Arthur Szlam, and Jason Weston. 2022a. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, and Shihang Wang. 2022b. Long time no see! open-domain conversation with long-term persona memory. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2639–2650, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213, Melbourne, Australia. Association for Computational Linguistics.
  35. 35.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278, Online. Association for Computational Linguistics.
  36. 36.Chujie Zheng, Sahand Sabour, Jiaxin Wen, Zheng Zhang, and Minlie Huang. 2023. Augesc: Dialogue augmentation with large language models for emotional support conversation.

Citation

MLA
Jang, J., et al. “Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13584–606, https://doi.org/10.18653/v1/2023.emnlp-main.838.
APA
Jang, J., Boo, M., & Kim, H. (2023). Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13584–13606. https://doi.org/10.18653/v1/2023.emnlp-main.838
Chicago
Jang, J., M. Boo, and H. Kim. 2023. “Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13584–606. https://doi.org/10.18653/v1/2023.emnlp-main.838.
Harvard
Jang, J., Boo, M. and Kim, H. (2023) “Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13584–13606. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.838.
Vancouver
1. Jang J, Boo M, Kim H (2023) Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13584–13606

BibTeX

@inproceedings{jang-etal-2023-conversation,
    title = "Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations",
    author = "Jang, Jihyoung  and
      Boo, Minseong  and
      Kim, Hyounghun",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.838/",
    doi = "10.18653/v1/2023.emnlp-main.838",
    pages = "13584--13606"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/