Evaluating Very Long-Term Conversational Memory of LLM Agents
Adyasha MaharanaDong-Ho LeeSergey TulyakovMohit BansalFrancesco BarbieriYuwei Fang
Introduces LoCoMo, a benchmark of human-verified, multi-modal conversations spanning dozens of sessions, revealing critical deficiencies in how modern long-context models and retrieval-augmented systems track long-range temporal and causal information compared to humans.
Existing research on conversational artificial intelligence has largely evaluated chatbots over short interactions, typically spanning fewer than five sessions and around one thousand tokens. However, real-world deployment requires systems to maintain consistent, empathetic, and coherent relationships over weeks or months. Modern large language models equipped with extended context windows or retrieval-augmented generation techniques are widely assumed to handle extended interactions, but their ability to track long-term conversational memory, causal timelines, and multimodal interactions has remained largely untested.
The article aims to evaluate how effectively state-of-the-art language models maintain very long-term conversational memory. It introduces a comprehensive evaluation benchmark to measure model performance in recalling past context, synthesizing temporal and causal event dynamics, and generating consistent multimodal dialogue over extended time horizons.
To conduct this evaluation, the researchers developed a hybrid machine-human pipeline to construct a benchmark dataset called LOCOMO. The pipeline used generative agents grounded in distinct personal backgrounds, chronologically ordered event graphs spanning six to twelve months, and image-sharing behaviors. Human annotators then edited approximately 15% of the dialogue turns and 19% of the images to resolve inconsistencies and ensure strict narrative alignment. The resulting dataset comprises 10 very long-term conversations averaging roughly 600 turns, 27 sessions, and over 16,000 tokens each. The authors evaluated base models, long-context models, and retrieval-augmented systems across three core tasks: five-category question answering, event graph summarization, and multimodal dialogue generation.
The findings show that current language models struggle substantially with very long-term memory. While long-context models and retrieval systems improve question-answering accuracy over base models by 12% to 20%, even the best-performing model (GPT-4-Turbo at an overall score of 51.6%) lags significantly behind human performance (87.9%), with a 41% deficit in temporal reasoning. Furthermore, long-context models suffer a severe vulnerability to adversarial questions, dropping by up to 65% compared to shorter-context baselines because large context windows easily mislead them into generating hallucinations and misattributing statements to the wrong speaker. In event summarization, models frequently miss causal links and confuse social cues like humor or sarcasm. In multimodal dialogue generation, retrieval-augmented generation using structured factual observations yielded the best results, though model relevance degraded as dialogue history lengthened.
These results demonstrate that expanding model context windows alone does not solve conversational memory. Long-context models can locate facts across broad contexts but fail to reason over them accurately, introducing operational risks such as hallucinations, incorrect user attribution, and misinterpretation of conversational nuance. For organizations deploying conversational systems, relying solely on broad context windows presents performance and reliability risks, whereas structuring historical interactions into factual observation databases offers a more robust near-term architecture.
Organizations developing conversational agents should avoid relying solely on long-context processing for extended interactions. Instead, teams should implement retrieval-augmented pipelines that distill past conversations into structured, factual observations rather than raw chat logs or summaries. In addition, systems deployed in customer-facing roles must include safeguards against adversarial hallucinations and speaker confusion, along with transparent disclosures regarding synthetic dialogue generation to mitigate user over-reliance.
The findings carry moderate limitations. The evaluation relies on a benchmark of ten human-edited, synthetically generated conversations and utilizes web-sourced imagery that lacks personal visual continuity across sessions. Additionally, standard automated metrics face inherent difficulty evaluating varied long-form model outputs. Consequently, readers should view these findings as an informative baseline rather than a definitive measure of human conversational behavior, and practitioners should conduct targeted pilot tests before deploying long-term conversational agents in high-stakes environments.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT establishes the foundational hierarchical memory and operating-system-inspired context paging architectures that long-term conversational memory evaluations build upon.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench introduces the multi-task evaluation methodology for long-context language model understanding that underpins extended conversational benchmarks.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This paper introduces persona-grounded multi-turn dialogue modeling, establishing the persona consistency framework evaluated in long-term conversational memory.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen provides the multi-agent conversational framework leveraged to generate and coordinate complex, multi-session agent interactions.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey formalizes the core architectural taxonomy for LLM-based autonomous agents, detailing the memory and profiling modules essential for sustaining multi-session dialogues.
- Paper: Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao et al. (2023). StreamingLLM addresses the fundamental attention and context-window degradation issues that occur when processing lengthy multi-round conversations.
- Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). LaMDA introduces core techniques for factual grounding and multi-turn open-domain dialogue quality that inform conversational agent benchmarks.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Synthesizes advanced agentic reasoning paradigms, expanding on long-term memory and dynamic multi-agent interaction challenges identified in benchmarks like LoCoMo.
- Paper: Recursive Language Models, Alex L. Zhang et al. (2025). Proposes programmatic, recursive prompt management to scale LLM effective context across tens of millions of tokens, offering a solution to context-window bottlenecks observed in long-term dialogues.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). Develops question-aware prompt compression and reordering to overcome the 'lost in the middle' and latency hurdles exposed in very long context evaluations.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). Extends fine-grained, instance-specific model evaluation methodologies across core capabilities including grounding and planning using evaluator LLMs.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Provides a comprehensive taxonomy and analysis of LLM-as-a-judge approaches for evaluating nuanced multi-turn conversational outputs and agent reliability.
