Evaluating Very Long-Term Conversational Memory of LLM Agents

Adyasha MaharanaDong-Ho LeeSergey TulyakovMohit BansalFrancesco BarbieriYuwei Fang

article2024ACL359 citations

Introduces LoCoMo, a benchmark of human-verified, multi-modal conversations spanning dozens of sessions, revealing critical deficiencies in how modern long-context models and retrieval-augmented systems track long-range temporal and causal information compared to humans.

Listen

Existing research on conversational artificial intelligence has largely evaluated chatbots over short interactions, typically spanning fewer than five sessions and around one thousand tokens. However, real-world deployment requires systems to maintain consistent, empathetic, and coherent relationships over weeks or months. Modern large language models equipped with extended context windows or retrieval-augmented generation techniques are widely assumed to handle extended interactions, but their ability to track long-term conversational memory, causal timelines, and multimodal interactions has remained largely untested.

The article aims to evaluate how effectively state-of-the-art language models maintain very long-term conversational memory. It introduces a comprehensive evaluation benchmark to measure model performance in recalling past context, synthesizing temporal and causal event dynamics, and generating consistent multimodal dialogue over extended time horizons.

To conduct this evaluation, the researchers developed a hybrid machine-human pipeline to construct a benchmark dataset called LOCOMO. The pipeline used generative agents grounded in distinct personal backgrounds, chronologically ordered event graphs spanning six to twelve months, and image-sharing behaviors. Human annotators then edited approximately 15% of the dialogue turns and 19% of the images to resolve inconsistencies and ensure strict narrative alignment. The resulting dataset comprises 10 very long-term conversations averaging roughly 600 turns, 27 sessions, and over 16,000 tokens each. The authors evaluated base models, long-context models, and retrieval-augmented systems across three core tasks: five-category question answering, event graph summarization, and multimodal dialogue generation.

The findings show that current language models struggle substantially with very long-term memory. While long-context models and retrieval systems improve question-answering accuracy over base models by 12% to 20%, even the best-performing model (GPT-4-Turbo at an overall score of 51.6%) lags significantly behind human performance (87.9%), with a 41% deficit in temporal reasoning. Furthermore, long-context models suffer a severe vulnerability to adversarial questions, dropping by up to 65% compared to shorter-context baselines because large context windows easily mislead them into generating hallucinations and misattributing statements to the wrong speaker. In event summarization, models frequently miss causal links and confuse social cues like humor or sarcasm. In multimodal dialogue generation, retrieval-augmented generation using structured factual observations yielded the best results, though model relevance degraded as dialogue history lengthened.

These results demonstrate that expanding model context windows alone does not solve conversational memory. Long-context models can locate facts across broad contexts but fail to reason over them accurately, introducing operational risks such as hallucinations, incorrect user attribution, and misinterpretation of conversational nuance. For organizations deploying conversational systems, relying solely on broad context windows presents performance and reliability risks, whereas structuring historical interactions into factual observation databases offers a more robust near-term architecture.

Organizations developing conversational agents should avoid relying solely on long-context processing for extended interactions. Instead, teams should implement retrieval-augmented pipelines that distill past conversations into structured, factual observations rather than raw chat logs or summaries. In addition, systems deployed in customer-facing roles must include safeguards against adversarial hallucinations and speaker confusion, along with transparent disclosures regarding synthetic dialogue generation to mitigate user over-reliance.

The findings carry moderate limitations. The evaluation relies on a benchmark of ten human-edited, synthetically generated conversations and utilizes web-sourced imagery that lacks personal visual continuity across sessions. Additionally, standard automated metrics face inherent difficulty evaluating varied long-form model outputs. Consequently, readers should view these findings as an informative baseline rather than a definitive measure of human conversational behavior, and practitioners should conduct targeted pilot tests before deploying long-term conversational agents in high-stakes environments.

arXiv: 2402.17753snap-research/locomo
Cover for Evaluating Very Long-Term Conversational Memory of LLM Agents

Abstract

Existing works on long-term open-domain dialogues focus on evaluating model responses within contexts spanning no more than five chat sessions. Despite advancements in long-context large language models (LLMs) and retrieval augmented generation (RAG) techniques, their efficacy in very long-term dialogues remains unexplored. To address this research gap, we introduce a machine-human pipeline to generate high-quality, very long-term dialogues by leveraging LLM-based agent architectures and grounding their dialogues on personas and temporal event graphs. Moreover, we equip each agent with the capability of sharing and reacting to images. The generated conversations are verified and edited by human annotators for long-range consistency and grounding to the event graphs. Using this pipeline, we collect LoCoMo, a dataset of very long-term conversations, each encompassing approx. 600 turns and 16K tokens on avg., over up to 32 sessions. Based on LoCoMo, we present a comprehensive evaluation benchmark to measure long-term memory in models, encompassing question answering, event summarization, and multi-modal dialogue generation tasks. Our experimental results indicate that LLMs exhibit challenges in understanding lengthy conversations and comprehending long-range temporal and causal dynamics within dialogues. Employing strategies like long-context LLMs or RAG can offer improvements but these models still substantially lag behind human performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Generative Pipeline for LOCOMO
  • 3.1 Persona
  • 3.2 Temporal Event Graph
  • 3.3 Virtual Agent Architecture
  • 3.4 Human Verification & Editing
  • 4 LOCOMO Evaluation Benchmark
  • 4.1 Question Answering Task
  • 4.2 Event Summarization Task
  • 4.3 Multi-Modal Dialogue Generation Task
  • 5 Experimental Setup
  • 6 Experimental Results
  • 6.1 Question Answering Task
  • 6.2 Event Summarization Task
  • 6.3 Multi-Modal Dialog Generation Task
  • 7 Conclusion
  • 8 Limitations
  • 9 Broader Impacts
  • References
  • Overview
  • A Generative Pipeline for LOCOMO
  • A.1 Persona
  • A.2 Temporal Event Graph
  • A.2.1 Virtual Agent Architecture
  • A.3 Human Filtering
  • B Dataset
  • B.1 Dataset Statistics
  • B.2 Dataset License
  • B.3 Annotator Details
  • C Experimental Setup
  • C.1 Baselines
  • C.2 Implementation Details
  • D Results
  • D.1 Event Summarization Task
  • D.2 Multimodal Dialog Generation Task

Knowls

  1. Knowl 1 — LOCOMO Dataset Specification and Statistics

    definition

    The LOCOMO (Long-term Conversational Memory) dataset is a benchmark designed to evaluate long-term conversational memory, reasoning, and multimodal grounding of large language model (LLM) agents across multi-session dialogues spanning several months.

    The dataset consists of 10 extensive multi-session conversations between pairs of virtual agents. Its core quantitative properties and benchmark subsets are as follows:

    • Conversation Scale: Each conversation contains an average of 27.227.2 sessions (ranging up to 3232 sessions), 588.2588.2 total turns (21.621.6 turns per session on average), and 16,618.116,618.1 tokens on average. An individual dialogue turn hkjh_k^j averages 29.829.8 tokens.
    • Temporal and Memory Grounding: Conversations average 19.219.2 tokens per atomic observation okjo_k^j, 132.4132.4 tokens per session summary wkw_k, and 35.835.8 ground-truth life events per conversation.
    • Multimodal Content: Conversations contain an average of 91.291.2 images retrieved and reacted to throughout the dialogue history.
    • Question Answering Subset: Comprises 1,9861,986 question-answer pairs categorized into single-hop retrieval (841841 questions, 42.3%42.3\%), multi-hop retrieval (282282 questions, 14.2%14.2\%), temporal reasoning (321321 questions, 16.1%16.1\%), open-domain knowledge (9696 questions, 4.8%4.8\%), and adversarial questions (446446 questions, 22.4%22.4\%).
    • Event Summarization Subset: Features reference event summaries averaging 1,042.71,042.7 tokens per conversation, derived from underlying temporal event graphs.
    • Licensing: Released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.
  2. Knowl 2 — Machine-Human Generative Pipeline for Long-Term Multimodal Dialogues

    model/method

    The LOCOMO generative pipeline produces multi-session, long-term conversations through a four-stage process combining LLM generative agents and human editing:

    1. Persona Expansion: Initial short persona statements pcp_c (4–5 sentences sourced from the MSC dataset) are expanded via gpt-3.5-turbo into comprehensive persona profiles pp containing demographics, objectives, daily habits, past experiences, and interpersonal relationships.
    2. Temporal Event Graph Generation: For each agent persona pp, an LLM (text-davinci-003) constructs a directed temporal event graph G={ei}i=1N\mathcal{G} = \{e_i\}_{i=1}^N (N≤25N \le 25 events across 6 to 12 months). Each event eie_i is assigned a calendar date tit_i and causal links l=(ei,ej)l = (e_i, e_j) specifying that past event eie_i caused subsequent event eje_j. The graph is generated iteratively starting from an initial seed of k=3k=3 independent events.
    3. Virtual Agent Architecture with Dual Memory and Multimodal Modules:
      • Short-Term Memory (HsH_s): After each session kk, an LLM generates a session summary wkw_k conditioned on the session transcript hkh_k and previous summary wk−1∈Hsw_{k-1} \in H_s.
      • Long-Term Memory (HlH_l): For every turn jj in session kk, the turn content hkjh_k^j is transformed into an objective, factual assertion (observation okjo_k^j) and stored in HlH_l alongside contributing turn IDs.
      • Response Generation: At session k+1k+1 on date tk+1st_{k+1}^s, agent LiL_i generates responses conditioned on wkw_k, relevant observations retrieved from HlH_l, ongoing session context hk+1h_{k+1}, persona pp, and inter-session events {e∈G∣tks<tie<tk+1s}\{e \in \mathcal{G} \mid t_k^s < t_i^e < t_{k+1}^s\}.
      • Image Sharing & Reaction: When sending an image, the agent generates a descriptive caption cc, extracts search keywords ww, queries the web via an image crawler, and sends the retrieved image. When receiving an image, the recipient captions it using BLIP-2 and generates a grounded textual reaction.
    4. Human Verification and Post-Editing: Annotators manually inspect and edit conversations to resolve long-range inconsistencies (editing ∼15%\sim 15\% of dialogue turns), align dialogue claims with event graphs G\mathcal{G}, and remove or replace irrelevant/incoherent images (substituting ∼19%\sim 19\% of images).
  3. Knowl 3 — Reasoning Categories and Evaluation Metrics for Long-Term Dialogue QA

    experimental setup

    The LOCOMO question answering benchmark evaluates an agent's memory retention and reasoning over long dialogue histories (1,9861,986 total QA instances). Questions are categorized into five distinct reasoning types:

    1. Single-Hop Retrieval (42.3%42.3\%): Requires recalling a localized piece of information stated within a single session.
    2. Multi-Hop Retrieval (14.2%14.2\%): Requires synthesizing distributed information across multiple disjoint sessions.
    3. Temporal Reasoning (16.1%16.1\%): Requires tracking calendar dates, relative time cues, time intervals, and chronological order across sessions.
    4. Open-Domain Knowledge (4.8%4.8\%): Requires combining facts disclosed by a speaker with external commonsense or world knowledge.
    5. Adversarial (22.4%22.4\%): Designed to mislead the agent into generating hallucinated responses about events that never occurred or attributing facts to the wrong speaker. The ground truth answer requires identifying the question as unanswerable.

    Evaluation Metrics:

    • Answer Prediction: Evaluated via token-level partial match F1 score between normalized model predictions and ground-truth reference strings, where ground-truth answers are extracted verbatim from the dialogue transcripts.
    • Retrieval Recall: For Retrieval-Augmented Generation (RAG) models, context retrieval quality is measured by Recall@kk (R@kR@k), defining whether the ground-truth turn IDs (or session summaries) containing the answer were retrieved within the top-kk documents.
  4. Knowl 4 — Temporal Event Graph Summarization via FactScore Adaptation

    model/method

    The event summarization task evaluates an agent's ability to extract, synthesize, and recount chronologically ordered life events from a lengthy dialogue history against the ground-truth temporal event graph G\mathcal{G}.

    To bypass the limitations of surface-level overlap metrics (BLEU, ROUGE) which fail to evaluate factual accuracy, the evaluation adapts FactScore by decomposing both the ground-truth event graph G\mathcal{G} and the model's generated summary into atomic factual statements:

    Precision=∣Atomic facts in generated summary that match G∣∣Total atomic facts in generated summary∣\text{Precision} = \frac{|\text{Atomic facts in generated summary that match } \mathcal{G}|}{|\text{Total atomic facts in generated summary}|}

    Recall=∣Atomic facts in G captured in generated summary∣∣Total atomic facts in G∣\text{Recall} = \frac{|\text{Atomic facts in } \mathcal{G} \text{ captured in generated summary}|}{|\text{Total atomic facts in } \mathcal{G}|}

    F1=2⋅Precision⋅RecallPrecision+Recall\text{F1} = \frac{2 \cdot \text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}

    For base models with limited context windows, an incremental summarization strategy is employed, wherein the model iteratively summarizes preceding sessions and uses the resulting summary as input context when summarizing subsequent sessions.

  5. Knowl 5 — Question Answering Performance of Base and Long-Context LLMs

    data/table

    Evaluating open-source base models and proprietary long-context LLMs across LOCOMO QA reasoning categories reveals significant performance gaps between LLMs and human capabilities, as well as a severe vulnerability of long-context models to adversarial questions.

    Category Model Context Single-Hop Multi-Hop Temporal Open Domain Adversarial Overall F1
    Human Human - 95.1 85.8 92.6 75.4 89.4 87.9
    Base Mistral-7B-Instruct-v0.2 8K 19.1 15.1 9.3 8.6 28.9 18.7
    Base Llama-2-70b-chat 4K 20.8 18.2 15.9 18.8 15.7 18.4
    Base Llama-3-70B-Instruct 4K 17.0 17.0 12.0 13.0 80.0 30.1
    Long Context gpt-3.5-turbo 4K 23.8 18.0 15.6 20.4 34.8 23.9
    Long Context gpt-3.5-turbo 8K 38.5 25.1 22.7 25.9 28.7 31.2
    Long Context gpt-3.5-turbo 12K 45.7 32.4 25.5 23.4 21.5 34.0
    Long Context gpt-3.5-turbo 16K 52.6 36.7 24.3 24.0 14.8 35.9
    Long Context gemini-1.0-pro 1M 62.4 35.3 34.2 19.0 5.2 39.1
    Long Context claude-3-sonnet 200K 70.7 38.1 26.9 52.2 2.5 42.8
    Long Context gpt-4-turbo 128K 72.3 51.5 51.4 38.5 15.7 51.6

    Key empirical findings from this comparison include:

    • Human vs. Machine Gap: The top-performing model (gpt-4-turbo, 51.651.6 overall F1) lags behind human performance (87.987.9 F1) by 36.336.3 absolute points, with the largest deficit in temporal reasoning (51.451.4 vs. 92.692.6, a 41.241.2-point gap).
    • Adversarial Fragility: Expanding the context window causes long-context LLMs to hallucinate answers on adversarial queries rather than recognizing them as unanswerable. For instance, gpt-3.5-turbo's adversarial F1 drops from 34.8%34.8\% at 4K4\text{K} to 14.8%14.8\% at 16K16\text{K}, while claude-3-sonnet and gemini-1.0-pro score only 2.5%2.5\% and 5.2%5.2\%, respectively, compared to 80.0%80.0\% for Llama-3-70B-Instruct operating on truncated context.
  6. Knowl 6 — Memory Granularity in Retrieval-Augmented Generation for Long-Term Dialogue QA

    data/table

    Evaluating Retrieval-Augmented Generation (RAG) using a DRAGON dense retriever and a gpt-3.5-turbo reader demonstrates that the structural representation of stored conversational memory strongly dictates answer quality and signal-to-noise ratio.

    Answer Prediction (F1) Recall Accuracy (R@k)
    Unit top-kk Single Multi Temp Open Adv All Single Multi Temp Open Adv All
    None - 29.9 23.3 17.5 29.5 12.8 22.4 - - - - - -
    Dialog 5 53.3 31.2 35.4 25.0 21.5 38.8 68.0 35.4 70.4 33.1 43.9 56.7
    Dialog 10 56.9 34.6 34.5 23.9 17.5 39.7 77.7 46.6 77.3 40.8 54.7 66.2
    Dialog 25 59.9 38.7 37.2 25.0 12.8 41.0 87.1 62.5 83.5 52.6 66.3 76.7
    Dialog 50 60.1 40.6 36.9 22.4 9.9 40.5 91.1 73.0 89.6 61.7 72.5 82.7
    Observation 5 54.3 36.3 40.7 26.5 32.5 43.3 67.1 41.4 73.1 35.1 37.7 56.2
    Observation 10 54.6 39.2 40.5 24.4 28.5 42.8 70.5 50.9 76.4 37.6 45.1 61.3
    Observation 25 54.7 41.2 38.8 25.8 24.7 42.1 74.7 60.6 80.3 49.0 53.5 67.5
    Observation 50 54.1 41.6 37.6 24.4 20.4 40.6 76.3 67.8 82.0 55.9 59.5 71.2
    Summary 2 32.8 22.4 32.9 19.0 25.3 29.0 63.0 33.0 57.7 31.5 65.9 65.9
    Summary 5 35.1 26.0 37.4 21.2 23.5 30.9 77.0 54.3 72.8 47.6 79.1 72.1
    Summary 10 36.0 29.9 37.5 22.2 24.0 32.0 88.7 72.7 84.5 67.3 88.8 84.7

    Key takeaways regarding retrieval memory formats:

    • Superiority of Atomic Observations: Storing conversational history as extracted atomic assertions (observations) yields the highest overall QA accuracy (43.343.3 F1 at top-55), outperforming raw dialogue turns (38.838.8 F1) and session summaries (30.930.9 F1).
    • Signal-to-Noise Degradation: Increasing the number of retrieved observations from k=5k=5 to k=50k=50 decreases overall answer F1 from 43.343.3 to 40.640.6 and adversarial F1 from 32.532.5 to 20.420.4, despite increasing retrieval recall from 56.2%56.2\% to 71.2%71.2\%.
    • Information Loss in Summaries: Although session summaries achieve the highest retrieval recall (84.7%84.7\% at top-1010), they yield poor answer prediction (32.032.0 F1) due to the loss of fine-grained conversational nuances during session summarization.
  7. Knowl 7 — Long-Context Event Summarization Performance and Error Taxonomy

    data/table

    Evaluating event summarization across base and long-context LLMs demonstrates the challenge of extracting chronological and causal life narratives from multi-session dialogues:

    ROUGE FactScore
    Category Model Context ROUGE-1 ROUGE-2 ROUGE-L Precision Recall FactScore F1
    Base Mistral-7B-Instruct-v0.2 8K 34.6 10.1 16.4 33.5 31.2 32.3
    Base Llama-3-70B-Instruct 4K 36.7 11.4 19.2 40.3 35.6 37.8
    Long Context gemini-1.0-pro 1M 37.6 13.4 21.1 46.7 42.1 44.2
    Long Context claude-3-sonnet 200K 35.1 12.6 21.3 45.6 40.8 43.1
    Long Context gpt-4-turbo 128K 41.2 13.8 21.6 51.9 46.5 48.9

    Error Taxonomy in LLM Event Summaries:

    1. Missing Information: Omitting critical event details or causal connections due to failure in linking distant dialogue turns across sessions.
    2. Hallucination: Padding summaries with invented claims or conflating attributes from separate events occurring within the same session.
    3. Misunderstanding of Dialogue Cues: Misinterpreting figurative language, humor, sarcasm, or hypotheticals as literal factual life events.
    4. Speaker Attribution Errors: Assigning an event, experience, or preference to the incorrect conversational participant.
    5. Saliency Errors: Extracting trivial conversational exchanges or phatic chit-chat as salient life events.
  8. Knowl 8 — Context-Augmented Multimodal Dialogue Generation Performance

    data/table

    Multimodal dialogue generation experiments using MiniGPT-5 (initialized from checkpoints fine-tuned on MMDialog) show the effect of conditioning response generation on different memory representations:

    Context Variant top-kk BLEU-1 / BLEU-2 Rouge-L MM-Relevance
    Base (Prior turns only) - 56.4 / 31.8 11.6 54.2
    + Summary 1 57.2 / 31.6 11.9 54.7
    + Summary 2 56.6 / 30.9 11.7 54.1
    + Summary 5 56.2 / 30.5 11.5 54.0
    + Observation 5 58.7 / 32.2 12.6 55.8
    + Observation 10 58.1 / 32.1 12.0 55.1
    + Observation 25 57.8 / 31.6 11.8 54.9

    Key results include:

    • Conditioning on retrieved atomic observations at top-55 achieves the highest text quality (58.758.7 BLEU-1, 12.612.6 Rouge-L) and multimodal alignment (55.855.8 MM-Relevance).
    • In vanilla MiniGPT-5 generation, MM-Relevance degrades monotonically as the length of prior dialogue history increases. Incorporating retrieval-augmented observation conditioning mitigates this degradation by grounding image and text generation in historical persona facts.
  9. Knowl 9 — Limitations of the LOCOMO Benchmark and Evaluation Setup

    limitation

    The LOCOMO dataset and evaluation benchmark exhibit five primary methodological limitations:

    1. Synthetic Generation Artefacts: Conversations are generated via synthetic LLM agent interactions before human post-editing; despite manual filtering of 15%15\% of turns and 19%19\% of images, the dialogues may not fully capture the naturalistic nuances, interruptions, and socio-emotional depth of real human online relationships.
    2. Visual Continuity Constraints: Images are retrieved from web search engines rather than generated as persistent personal photo collections. Consequently, they lack temporal visual consistency regarding character appearances, environments, pets, and personal objects.
    3. Language Scope: The generative framework, prompt pipelines, and evaluation benchmarks are implemented solely in English.
    4. Proprietary LLM API Dependency: The data generation pipeline relies on closed-source, commercial LLM APIs (gpt-3.5-turbo, text-davinci-003), posing reproducibility challenges as proprietary API model versions evolve or are deprecated.
    5. Long-Form NLG Evaluation Sensitivity: Automatic partial-match F1 evaluation can be sensitive to verbosity and paraphrasing variations inherent to LLM generation outputs.

Coverage note — None was omitted; all contributed benchmark specifications, generative pipeline components, evaluation tasks, baseline results across Base/Long-Context/RAG LLMs, error taxonomies, and stated limitations are fully represented.

References

  1. 1.Jaewoo Ahn, Yeda Song, Sangdoo Yun, and Gunhee Kim. 2023. Mpchat: Towards multimodal persona-grounded conversation. arXiv preprint arXiv:2305.17388.
  2. 2.Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021. Open-domain question answering goes conversational via question rewriting. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 520–534.
  3. 3.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433.
  4. 4.Jan Assmann and John Czaplicka. 1995. Collective memory and cultural identity. New german critique, (65):125–133.
  5. 5.Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew R Gormley. 2023. Unlimiformer: Longrange transformers with unlimited length input. arXiv preprint arXiv:2305.01625.
  6. 6.Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2023. Booookscore: A systematic exploration of book-length summarization in the era of llms. arXiv preprint arXiv:2310.00785.
  7. 7.Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2022. Summscreen: A dataset for abstractive screenplay summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8602–8615.
  8. 8.Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2023. Longlora: Efficient fine-tuning of long-context large language models. arXiv preprint arXiv:2309.12307.
  9. 9.Alan Cooper. 1999. The inmates are running the asylum. Springer.
  10. 10.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359.
  11. 11.Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017. Visual dialog. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 326–335.
  12. 12.Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin. 2022. Mmdialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation. arXiv preprint arXiv:2211.05719.
  13. 13.Silin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, and Antoine Bosselut. 2023a. Peacok: Persona commonsense knowledge for consistent and engaging narratives. arXiv preprint arXiv:2305.02364.
  14. 14.Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023b. Enabling large language models to generate text with citations. arXiv preprint arXiv:2305.14627.
  15. 15.Sarik Ghazarian, Nuan Wen, Aram Galstyan, and Nanyun Peng. 2022. Deam: Dialogue coherence evaluation using amr-based semantic manipulations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 771–785.
  16. 16.Xingwei He, Zhenghao Lin, Yeyun Gong, Alex Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023. Annollm: Making large language models to be better crowdsourced annotators. arXiv preprint arXiv:2303.16854.
  17. 17.William Hirst and Gerald Echterhoff. 2012. Remembering in conversations: The social sharing and reshaping of memories. Annual review of psychology, 63:55–79.
  18. 18.William Hirst and David Manier. 2008. Towards a psychology of collective memory. Memory, 16(3):183–200.
  19. 19.William Hirst, Jeremy K Yamashiro, and Alin Coman. 2018. Collective memory from a psychological perspective. Trends in cognitive sciences, 22(5):438–451.
  20. 20.Pegah Jandaghi, XiangHai Sheng, Xinyi Bai, Jay Pujara, and Hakim Sidahmed. 2023. Faithful persona-based conversational dataset generation with large language models. arXiv preprint arXiv:2312.10007.
  21. 21.Jihyoung Jang, Minseong Boo, and Hyounghun Kim. 2023a. Conversation chronicles: Towards diverse temporal and relational dynamics in multi-session conversations. arXiv preprint arXiv:2310.13420.
  22. 22.Jihyoung Jang, Minseong Boo, and Hyounghun Kim. 2023b. Conversation chronicles: Towards diverse temporal and relational dynamics in multi-session conversations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13584–13606, Singapore. Association for Computational Linguistics.
  23. 23.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  24. 24.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, and Yejin Choi. 2023. SODA: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12930–12949, Singapore. Association for Computational Linguistics.
  25. 25.Satwik Kottur, José MF Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2019. Clevr-dialog: A diagnostic dataset for multi-round reasoning in visual dialog. arXiv preprint arXiv:1903.03166.
  26. 26.Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. 2023. Longeval: Guidelines for human evaluation of faithfulness in long-form summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 1642–1661.
  27. 27.Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. 2022. Booksum: A collection of datasets for long-form narrative summarization. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6536–6558.
  28. 28.Dong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White, and Sujay Jauhar. 2023a. Making large language models better data creators. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15349–15360, Singapore. Association for Computational Linguistics.
  29. 29.Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, and Kangwook Lee. 2023b. Prompted llms as chatbot modules for long open-domain conversation. arXiv preprint arXiv:2305.04533.
  30. 30.Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim, Jonghwan Hyeon, and Ho-Jin Choi. 2023c. Dialogcc: An automated pipeline for creating high-quality multimodal dialogue datasets. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
  31. 31.Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. 2023a. Loogle: Can long-context language models understand long contexts? arXiv preprint arXiv:2311.04939.
  32. 32.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pretraining with frozen image encoders and large language models. In International Conference on Machine Learning.
  33. 33.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986–995.
  34. 34.Xinnian Liang, Bing Wang, Hui Huang, Shuangzhi Wu, Peihao Wu, Lu Lu, Zejun Ma, and Zhoujun Li. 2023. Unleashing infinite-length input capacity for largescale language models with self-controlled memory system. arXiv preprint arXiv:2304.13343.
  35. 35.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  36. 36.Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. How to train your dragon: Diverse augmentation towards generalizable dense retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6385–6400, Singapore. Association for Computational Linguistics.
  37. 37.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  38. 38.Nelson F Liu, Tianyi Zhang, and Percy Liang. 2023b. Evaluating verifiability in generative search engines. arXiv preprint arXiv:2304.09848.
  39. 39.Junru Lu, Siyu An, Mingbao Lin, Gabriele Pergola, Yulan He, Di Yin, Xing Sun, and Yunsheng Wu. 2023. Memochat: Tuning llms to use memos for consistent long-range open-domain conversation. arXiv preprint arXiv:2308.08239.
  40. 40.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822, Toronto, Canada. Association for Computational Linguistics.
  41. 41.Yuxian Meng, Shuhe Wang, Qinghong Han, Xiaofei Sun, Fei Wu, Rui Yan, and Jiwei Li. 2020. Openvidial: A large-scale, open-domain dialogue dataset with visual contexts. arXiv preprint arXiv:2012.15015.
  42. 42.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251.
  43. 43.Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios P Spithourakis, and Lucy Vanderwende. 2017. Imagegrounded conversations: Multimodal context for natural question and response generation. arXiv preprint arXiv:1701.08251.
  44. 44.Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021. I like fish, especially dolphins: Addressing contradictions in dialogue modeling. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1699–1713.
  45. 45.Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  46. 46.Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. arXiv preprint arXiv:2304.03442.
  47. 47.John Pruitt and Jonathan Grudin. 2003. Personas: practice and theory. In Proceedings of the 2003 conference on Designing for user experiences, pages 1–15.
  48. 48.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. arXiv preprint arXiv:2302.00083.
  49. 49.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. Replug: Retrievalaugmented black-box language models. arXiv preprint arXiv:2301.12652.
  50. 50.Michael Shum, Stephan Zheng, Wojciech Kryściński, Caiming Xiong, and Richard Socher. 2019. Sketchfill-ar: A persona-grounded chit-chat generation framework. arXiv preprint arXiv:1910.13008.
  51. 51.Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston. 2018. Image chat: Engaging grounded conversations. arXiv preprint arXiv:1811.00945.
  52. 52.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3784–3803.
  53. 53.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  54. 54.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  55. 55.Yuqing Wang and Yun Zhao. 2023. Tram: Benchmarking temporal reasoning for large language models. arXiv preprint arXiv:2310.00835.
  56. 56.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3731–3741.
  57. 57.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  58. 58.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453.
  59. 59.Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023. A critical evaluation of evaluations for long-form question answering. arXiv preprint arXiv:2305.18201.
  60. 60.Jing Xu, Arthur Szlam, and Jason Weston. 2022. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197.
  61. 61.Xiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song, Hao Zhang, and Jindong Chen. 2021. Photochat: A human-human dialogue dataset with photo sharing behavior for joint image-text modeling. arXiv preprint arXiv:2108.01453.
  62. 62.Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021a. Dynaeval: Unifying turn and dialogue level evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5676–5689.
  63. 63.Chen Zhang, Luis Fernando D’Haro, Qiquan Zhang, Thomas Friedrichs, and Haizhou Li. 2022. Finedeval: Fine-grained automatic dialogue-level evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3336–3355.
  64. 64.Qiang Zhang, Jason Naradowsky, and Yusuke Miyao. 2023. Mind the gap between conversations for improved long-term dialogue generation. arXiv preprint arXiv:2310.15415.
  65. 65.Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, and Mohit Bansal. 2021b. Emailsum: Abstractive email thread summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6895–6909.
  66. 66.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  67. 67.Kaizhi Zheng, Xuehai He, and Xin Eric Wang. 2023. Minigpt-5: Interleaved vision-and-language generation via generative vokens. arXiv preprint arXiv:2310.02239.
  68. 68.Yinhe Zheng, Guanyi Chen, Xin Liu, and Jian Sun. 2021. Mmchat: Multi-modal chat dataset on social media. arXiv preprint arXiv:2108.07154.
  69. 69.Wanjun Zhong, Lianghong Guo, Qiqi Gao, and Yanlin Wang. 2023. Memorybank: Enhancing large language models with long-term memory. arXiv preprint arXiv:2305.10250.
  70. 70.Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. The design and implementation of xiaoice, an empathetic social chatbot. Computational Linguistics, 46(1):53–93.

Citation

MLA
Maharana, A., et al. “Evaluating Very Long-Term Conversational Memory of LLM Agents”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13851–70, https://doi.org/10.18653/v1/2024.acl-long.747.
APA
Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F., & Fang, Y. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13851–13870. https://doi.org/10.18653/v1/2024.acl-long.747
Chicago
Maharana, A., D.-H. Lee, S. Tulyakov, M. Bansal, F. Barbieri, and Y. Fang. 2024. “Evaluating Very Long-Term Conversational Memory of LLM Agents”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13851–70. https://doi.org/10.18653/v1/2024.acl-long.747.
Harvard
Maharana, A. et al. (2024) “Evaluating Very Long-Term Conversational Memory of LLM Agents”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13851–13870. Available at: https://doi.org/10.18653/v1/2024.acl-long.747.
Vancouver
1. Maharana A, Lee D-H, Tulyakov S, Bansal M, Barbieri F, Fang Y (2024) Evaluating Very Long-Term Conversational Memory of LLM Agents. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13851–13870

BibTeX

@inproceedings{maharana-etal-2024-evaluating,
    title = "Evaluating Very Long-Term Conversational Memory of {LLM} Agents",
    author = "Maharana, Adyasha  and
      Lee, Dong-Ho  and
      Tulyakov, Sergey  and
      Bansal, Mohit  and
      Barbieri, Francesco  and
      Fang, Yuwei",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.747/",
    doi = "10.18653/v1/2024.acl-long.747",
    pages = "13851--13870"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/