GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations
Jingbo YangKwei-Herng LaiXiaowen WangShiyu ChangYaar HarariEvgeniy Gabrilovich
Introduces GroupMemBench, a benchmark for evaluating LLM agent memory in multi-party conversations, revealing that current memory systems struggle with speaker-grounded context and collapse to an average accuracy of only 46%.
Artificial intelligence agents increasingly operate as workplace collaborators and assistants within shared channels and multi-user threads. In these settings, an agent's utility depends heavily on its long-term memory system to extract, retain, and recall information across ongoing discussions. However, existing conversational memory architectures and evaluation benchmarks were designed almost exclusively for one-on-one, single-user interactions. This dyadic design fails to account for critical group dynamics, such as threaded debates, speaker-grounded belief tracking (knowing who said what to whom), and role-specific lexical shifts where different professionals use different terminology for the same concepts.
The article introduces GroupMemBench, an evaluation benchmark designed to assess how well artificial intelligence memory systems handle multi-party conversations. Its main objective is to evaluate whether current memory architectures can accurately condition extraction and retrieval on specific speaker identities, maintain concurrent and evolving beliefs across multiple users, and resolve audience-adapted language in team environments.
To construct the benchmark, the authors developed a graph-grounded synthesis pipeline that generated 120,000 multi-party messages across four workplace domains: Technology, Finance, Healthcare, and Manufacturing. The conversation generation enforced controllable reply structures, distinct user personas, and targeted communication dynamics. To create challenging evaluation queries, an adversarial generation framework produced questions across six categories—multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention—only accepting queries that defeated baseline retrieval systems through an iterative refinement loop.
The evaluation revealed substantial performance gaps in existing memory systems. First, leading agent memory systems experienced severe performance degradation in group settings, with the top-performing system achieving only 46.0% average accuracy. Second, performance dropped sharply on multi-party core tasks: knowledge update accuracy fell to 27.1%, and resolving role-specific term ambiguity dropped to 37.7%. Third, a simple BM25 keyword-retrieval baseline matched or outperformed four out of five specialized agent memory systems, achieving 43.2% average accuracy at virtually zero ingestion cost. Detailed error analysis demonstrated that 41% to 79% of total errors stemmed from retrieval bottlenecks rather than reasoning failures, confirming that memory ingestion pipelines discard essential speaker identities and structural thread context before the model can reason over them.
These findings indicate that existing memory ingestion mechanisms actively degrade performance in multi-user environments by flattening conversational hierarchies and stripping away speaker attribution. Organizations deploying conversational agents in collaborative spaces risk providing incorrect, stale, or conflicting information if they rely on architectures tuned for single-user settings. Furthermore, complex extraction and graph-building ingestion pipelines incurred costs up to $51 per domain and large storage footprints without delivering commensurate performance improvements over basic text retrieval.
Decision-makers and system architects should avoid deploying current dyadic agent memory mechanisms directly into multi-user team channels without structural modifications. Instead, development teams should prioritize re-engineering memory ingestion to preserve conversational topology and treat user identities as first-class indexing attributes. For near-term applications, engineering teams can adopt hybrid or raw text-retrieval baselines that retain original conversational contexts at substantially lower computational cost. Further research should focus on multi-user belief tracking and bridging role-specific vocabulary gaps before deploying autonomous memory agents in mission-critical group operations.
The findings are subject to several boundary conditions, as the benchmark was evaluated exclusively on English-language, text-only workplace simulations without images, attachments, or voice notes. Nevertheless, confidence in the diagnostic conclusions remains high, supported by extensive cross-domain testing and human-annotated validation of the evaluation protocols.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Introduces the foundational benchmark evaluating long-term conversational memory and temporal reasoning in dyadic LLM interactions, which GroupMemBench extends to multi-party scenarios.
- Paper: LM Agents for Coordinating Multi-User Information Gathering, Harsh Jhamtani et al. (2025). Examines how language model agents coordinate and gather fragmented information across multiple users in workplace environments, providing direct motivation for group memory evaluation.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). Establishes the standard tiered memory and paging architecture for conversational agents that GroupMemBench tests under complex multi-user dynamics.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Provides the causal evaluation framework for Theory-of-Mind and belief tracking in LLMs, which underpins the speaker-grounded belief evaluation in GroupMemBench.
- Paper: Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker, Melanie Sclar et al. (2023). Develops graph-based multi-character belief tracking for language models, providing key concepts for speaker-grounded memory modeling across multiple dialogue participants.
- Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). Formulates multi-hop reasoning over dynamic knowledge updates, establishing the core problem of propagating updated facts that GroupMemBench evaluates in multi-party conversations.
- Paper: ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation, Ziyi Liu et al. (2025). Investigates socio-cognitive modeling and multi-party conversational dynamics, laying the groundwork for evaluating agent interaction within group contexts.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). Proposes iterative retrieval and deep search memory architectures for complex long-context dialogues, representing the class of agent memory systems benchmarked in GroupMemBench.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). Proposes an iterative graph memory reconstruction method that directly addresses the retrieval collapse and structural information loss identified in GroupMemBench.
- Paper: Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents, Theo Rusu et al. (2026). Investigates graph-structured conversational representations and selective forgetting mechanisms for long-term agent memory across complex interactions.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). Develops a multi-tiered cognitive memory architecture incorporating semantic knowledge graphs and offline consolidation to address multi-session streaming interactions.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). Provides an execution-based evaluation environment isolating dynamic memory management and compression policies across extended multi-turn interactions.
- Paper: $δ$-mem: Efficient Online Memory for Large Language Models, Jingdi Lei et al. (2026). Presents an efficient online state-compression memory mechanism to prevent context degradation and maintain factual accuracy over extended conversations.
- Paper: Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models, Harshita Chopra et al. (2026). Introduces prospection-guided retrieval to overcome the limitations of standard lexical and semantic similarity matching in long-term conversational memory.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). Explores parametric memory models trained on synthesized cross-document reflections to enable structured, multi-step fact retrieval.
- Paper: Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction, Ryo Kamoi et al. (2026). Evaluates LLM capabilities in modeling complex, realistic multi-party conversation behaviors, including misunderstandings and inconsistent social dynamics.
