GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Jingbo YangKwei-Herng LaiXiaowen WangShiyu ChangYaar HarariEvgeniy Gabrilovich

article2026arXiv13 citations

Introduces GroupMemBench, a benchmark for evaluating LLM agent memory in multi-party conversations, revealing that current memory systems struggle with speaker-grounded context and collapse to an average accuracy of only 46%.

Listen

Artificial intelligence agents increasingly operate as workplace collaborators and assistants within shared channels and multi-user threads. In these settings, an agent's utility depends heavily on its long-term memory system to extract, retain, and recall information across ongoing discussions. However, existing conversational memory architectures and evaluation benchmarks were designed almost exclusively for one-on-one, single-user interactions. This dyadic design fails to account for critical group dynamics, such as threaded debates, speaker-grounded belief tracking (knowing who said what to whom), and role-specific lexical shifts where different professionals use different terminology for the same concepts.

The article introduces GroupMemBench, an evaluation benchmark designed to assess how well artificial intelligence memory systems handle multi-party conversations. Its main objective is to evaluate whether current memory architectures can accurately condition extraction and retrieval on specific speaker identities, maintain concurrent and evolving beliefs across multiple users, and resolve audience-adapted language in team environments.

To construct the benchmark, the authors developed a graph-grounded synthesis pipeline that generated 120,000 multi-party messages across four workplace domains: Technology, Finance, Healthcare, and Manufacturing. The conversation generation enforced controllable reply structures, distinct user personas, and targeted communication dynamics. To create challenging evaluation queries, an adversarial generation framework produced questions across six categories—multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention—only accepting queries that defeated baseline retrieval systems through an iterative refinement loop.

The evaluation revealed substantial performance gaps in existing memory systems. First, leading agent memory systems experienced severe performance degradation in group settings, with the top-performing system achieving only 46.0% average accuracy. Second, performance dropped sharply on multi-party core tasks: knowledge update accuracy fell to 27.1%, and resolving role-specific term ambiguity dropped to 37.7%. Third, a simple BM25 keyword-retrieval baseline matched or outperformed four out of five specialized agent memory systems, achieving 43.2% average accuracy at virtually zero ingestion cost. Detailed error analysis demonstrated that 41% to 79% of total errors stemmed from retrieval bottlenecks rather than reasoning failures, confirming that memory ingestion pipelines discard essential speaker identities and structural thread context before the model can reason over them.

These findings indicate that existing memory ingestion mechanisms actively degrade performance in multi-user environments by flattening conversational hierarchies and stripping away speaker attribution. Organizations deploying conversational agents in collaborative spaces risk providing incorrect, stale, or conflicting information if they rely on architectures tuned for single-user settings. Furthermore, complex extraction and graph-building ingestion pipelines incurred costs up to $51 per domain and large storage footprints without delivering commensurate performance improvements over basic text retrieval.

Decision-makers and system architects should avoid deploying current dyadic agent memory mechanisms directly into multi-user team channels without structural modifications. Instead, development teams should prioritize re-engineering memory ingestion to preserve conversational topology and treat user identities as first-class indexing attributes. For near-term applications, engineering teams can adopt hybrid or raw text-retrieval baselines that retain original conversational contexts at substantially lower computational cost. Further research should focus on multi-user belief tracking and bridging role-specific vocabulary gaps before deploying autonomous memory agents in mission-critical group operations.

The findings are subject to several boundary conditions, as the benchmark was evaluated exclusively on English-language, text-only workplace simulations without images, attachments, or voice notes. Nevertheless, confidence in the diagnostic conclusions remains high, supported by extensive cross-domain testing and human-annotated validation of the evaluation protocols.

arXiv: 2605.14498
Cover for GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Abstract

Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems that extract, retrieve, and apply information across long-running conversations. However, both existing memory systems and benchmarks are built around the dyadic, single-user setup, even though real deployments routinely span groups and channels with multiple users interacting with the agent and with each other. This mismatch leaves three properties of group memory unmeasured: (i) group dynamics that go beyond concatenated one-on-one chats, (ii) speaker-grounded belief tracking, where the per-user memory modeling is needed, and (iii) audience-adapted language, where Theory-of-Mind shifts produce role-specific vocabulary. We introduce GroupMemBench, a benchmark that exposes all three. A graph-grounded synthesis pipeline produces multi-party conversations with controllable reply structure and conditions each message on per-user personas and target audiences. An adversarial query pipeline then binds every question to a specific asker across six categories, spanning multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention, and iteratively searches challenging, realistic queries that reflect comprehensive memory capability. Benchmarking leading memory systems exposes a sharp collapse: the strongest one reaches only 46.0% average accuracy, with knowledge update at 27.1% and term ambiguity at 37.7%, while a simple BM25 baseline matches or exceeds most agent memory systems. This indicates current memory ingestion erases the structural and lexical features group memory depends on, leaving multi-user memory far from solved.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 GroupMemBench
  • 3.1 Problem Formulation
  • 3.2 Synthesizing Multi-Party Conversations
  • 3.3 Adversarial Query Construction
  • 3.3.1 Shared Generation Pipeline
  • 3.3.2 Dimensions and Type-Specific Strategies
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Performance-Efficiency Trade-off
  • 4.4 Failure Analysis
  • 5 Conclusion
  • References
  • A Graph Schema
  • B Path Sampling and Graph Walks
  • C Prompt Templates
  • D Hyperparameters and Realism Knobs
  • E Evaluation Rubrics for Synthetic Conversation
  • F Dataset Statistics
  • G Per-Domain Results
  • H Per-Domain Ingestion Cost and Storage
  • I LLM Judge Configuration and Reliability
  • J Worked example: information loss during memory ingestion
  • K Limitations
  • L Broader Impacts

Knowls

  1. Knowl 1 — Group memory requires structure, speaker grounding, and audience adaptation

    definition

    GroupMemBench treats group-memory evaluation as testing three distinct capabilities. Group dynamics means preserving relationships among messages—such as threaded replies, multi-user exchanges, and cross-topic shifts—rather than flattening a conversation into a turn sequence or independent one-to-one chats. Speaker-grounded belief tracking means retaining who expressed a fact or preference, since the correct update or answer can depend on the speaker and the user asking. Audience-adapted language means handling role-dependent ways of describing the same underlying information: a speaker may use specialized vocabulary with one audience and more general wording with another. The benchmark is designed to test these properties jointly in multi-party workplace conversations.

  2. Knowl 2 — Memory question answering is explicitly conditioned on the asking user

    definition

    Let U={u1,…,uK}U=\{u_1,\ldots,u_K\} be a group of KK human users and let uau_a denote the memory-augmented agent. The chronological conversation is C=(m1,…,mN)C=(m_1,\ldots,m_N), with each message mi=(ai,ti,ci)m_i=(a_i,t_i,c_i) consisting of author ai∈U∪{ua}a_i\in U\cup\{u_a\}, timestamp tit_i, and text cic_i. Starting from empty memory M0M_0, a memory system updates its state after each message according to Mi=fupdate(Mi−1,mi)M_i=f_{\mathrm{update}}(M_{i-1},m_i); the update function is left unspecified so it can represent different memory implementations. After reading the prefix C1:TC_{1:T} and forming MTM_T, the system receives a question qq from a specified user uq∈Uu_q\in U and returns a=fans(q,uq,MT)a=f_{\mathrm{ans}}(q,u_q,M_T). The answer is added to the shared conversation as an agent message and affects subsequent memory and turns. Unlike the dyadic case, the asker is a meaningful input: the same question can have different correct answers depending on that user's identity, role, history, or interpretation of vocabulary.

  3. Knowl 3 — Graph-grounded synthesis generates structured multi-party conversations

    model/method

    GroupMemBench separates conversation structure from message wording. A directed graph represents domains or projects, topics, time-bounded phases, users, and generated messages. Its relations encode domain-to-topic and topic-to-phase hierarchies, user and phase membership, shared-category links between related projects, message authorship, and reply-to links. Phase nodes carry a status goal and target date; user profiles encode role, tone, style, expertise, and prior messages, with distinct tone–style combinations assigned to different users.

    At each generation step, graph traversal selects the author, recipients, phase, and conversational context. The three path types are channel posts, threaded replies, and cross-project role-to-role bridges, sampled in proportions 0.200.20, 0.600.60, and 0.200.20, respectively. Posts favor phases with few or no posts; reply selection favors active phases and unanswered posts, uses preferential attachment with recency decay, and prefers replying to posts over replies. Cross-project paths connect projects sharing a category and favor a recipient with the source user's role. The language model then generates a message conditioned on the selected path, persona, phase, and recent situational history; GPT-5 is used for this generation. Context includes up to 50 recent phase messages, with three snippets available from older history, and a phase ledger tracking decided issues, open questions, and pending commitments. Six distinct sub-issues are initially specified per phase, and the ledger is refreshed every eight new messages.

    The generation policy adds realism perturbations: about 5% of messages contain a subtle factual error, stale reference, or drift; 25% of replies disagree substantively or conditionally; 20% include stylistic imperfections; and eligible phases may have a committed decision reversed, with the reversal recorded as evaluation metadata. The result is a corpus with explicit thread structure, speaker identities, timestamps, and persona-conditioned wording rather than a flat collection of synthetic turns.

  4. Knowl 4 — Adversarial query construction accepts questions that defeat retrieval

    model/method

    Each benchmark question is built from a multi-party conversation using a shared four-stage process. First, a lightweight filter proposes messages containing target entities, and an LLM classifier checks that the evidence supports a concrete answer. Second, a question-proposer LLM drafts a natural-language question from the anchor message and target answer. Third, a retrieval-based solver attempts the question and an LLM judge assesses its answer: if the solver succeeds, the question is iteratively refined to make it harder; if the solver fails, the question is accepted. Finally, the benchmark retains the question, gold answer, querying user ID, and evidence trace for each refinement round. Thus difficulty is operationalized by failure of a competent retrieval baseline, not merely asserted by the question writers.

  5. Knowl 5 — Six query categories isolate different group-memory capabilities

    definition

    GroupMemBench uses six question types. Multi-hop reasoning requires linking evidence across distant messages or threads, often through an indirect pivot message. Knowledge update tests whether the system returns the latest decision or consensus rather than an obsolete state, and can ask about the current value, what prompted a change, or the earlier value. Term ambiguity phrases a question in the asker's role-specific vocabulary while the supporting evidence uses another speaker's wording, testing whether the system links distinct expressions for the same referent. User-implicit reasoning uses first-person references that must be resolved to the identified asker. Temporal reasoning requires reasoning over message timestamps and relative time expressions. Abstention asks about information absent from the conversation and tests whether the system refrains from inventing an answer.

  6. Knowl 6 — The benchmark contains four matched-size domains and graph-guided conversations approach a real-chat quality reference

    data/table

    The synthesized corpus covers Technology, Finance, Healthcare, and Manufacturing, with 30,000 messages in each domain (120,000 total). The matched message counts control for corpus-size differences across domains, while the range of participant and role counts varies organizational complexity. The graph-guided generator was also assessed with GPT-5 as judge on six conversation-quality dimensions—naturalness, coherence, diversity, contextual relevance, momentum, and engagingness—using scores on a 0–5 scale averaged over 10 seeds at fixed temperature. Across most dimensions, graph-guided synthesis tracked an upper-bound reference made from authentic group-chat logs collected over a one-month horizon and substantially outperformed a single-prompt generation baseline.

    Statistic Technology Finance Healthcare Manufacturing
    Projects 7 6 10 10
    Topics 35 30 30 30
    Phases 74 66 79 76
    Participants 18 12 6 9
    Roles 6 7 4 5
    Posts 6,599 6,764 7,043 6,816
    Replies 19,897 19,722 20,099 20,048
    Cross-project replies 3,504 3,514 2,858 3,136
    Noise messages 1,503 1,535 1,510 1,542
    Decision points 20,788 19,920 10,301 7,681
    Decision changes 74 66 113 99
    Total messages 30,000 30,000 30,000 30,000
  7. Knowl 7 — Memory systems achieve only 46.01% average accuracy, with major weaknesses on updates and ambiguity

    empirical result

    Accuracy is GPT-5-judge question-answering accuracy against reference answers, aggregated across the four domains. During ingestion, systems without their own pretrained checkpoint use gpt-4o-mini; query answering uses GPT-5, and dense retrieval uses text-embedding-3-large. BM25 indexes the raw conversation without LLM or embedding-based ingestion. The values below are percentages; the final column is the reported average.

    Method Multi-Hop Update Ambiguity Implicit Temporal Abstention Average
    BM25 40.11 25.23 14.15 40.82 54.94 77.98 43.22
    text-embedding-3-large 36.26 23.36 21.70 46.94 32.72 75.23 38.04
    GraphRAG 12.09 14.02 19.81 14.29 5.56 66.97 20.56
    Mem0 21.98 4.67 11.32 20.41 16.67 82.57 25.73
    MemGPT 22.53 17.76 20.75 28.57 12.42 77.98 28.15
    A-Mem 35.16 22.43 26.42 46.94 23.46 67.89 35.10
    HippoRAG 39.56 27.10 30.19 42.86 29.63 75.23 39.72
    Hindsight 42.31 17.76 37.74 40.82 54.94 77.06 46.01

    Hindsight has the highest overall average, but the best score is still 46.01%. No system exceeds 28% on Knowledge Update, and no system reaches 38% on Term Ambiguity. BM25's 43.22% average matches or exceeds four of the five agent-memory systems, despite not using semantic memory transformations during ingestion.

  8. Knowl 8 — Most errors arise before reasoning because retrieval loses group structure or speaker context

    empirical result

    The failure analysis classifies non-abstention errors as retrieval failures when the gold supporting message is not retrieved, and reasoning failures when it is retrieved but the answer is wrong. Across systems, retrieval failures account for 41–79% of questions, compared with 7–24% reasoning failures; the analysis uses 185 non-abstention questions per baseline per domain. Answer accuracy generally rises with retrieval recall, indicating that much of the multi-hop and speaker-grounding deficit originates during memory ingestion, before the answering model sees evidence. A worked user-implicit example illustrates the mechanism: compressed topic-only memories can drop the request and its speaker, similar-topic messages from another user can shadow the asker's message, and retrieval of several users' requests can lead to an over-broad union of their answers.

    Term Ambiguity is the residual weakness after controlling for retrieval: conditional on retrieving the gold evidence, accuracy remains below 40% in nearly every evaluated cell, whereas Temporal Reasoning is uniformly high when its evidence is retrieved. BM25, which provides raw message text, has the highest mean conditional accuracy in both Technology and Finance. These results indicate that memory rewriting can impair reasoning over retrieved evidence, and that role-specific lexical shifts challenge the representation as well as the retrieval step.

  9. Knowl 9 — Higher ingestion cost does not reliably buy accuracy, and storage size does not predict performance

    empirical result

    Ingestion cost sums LLM and embedding calls and excludes query-time calls; storage is the measured on-disk footprint. Across per-domain runs, ingestion cost ranges from 0.31to0.31 to 51.04 and storage from 0.69 to 8.63 GB. The following values are unweighted means across the four domains for the six systems in the cost–storage comparison.

    Method Average ingestion cost (USD) Average storage (GB)
    Mem0 18.41 0.99
    MemGPT 0.43 0.69
    HippoRAG 8.64 5.76
    A-Mem 26.13 6.60
    Hindsight 32.10 4.25
    GraphRAG 12.51 3.00

    HippoRAG lies on the cost–accuracy Pareto front in all four domains, while A-Mem, Mem0, and GraphRAG are dominated in three or four domains. GraphRAG costs more than HippoRAG yet trails it by 14–25 absolute accuracy points in the domain comparisons. Storage does not track accuracy: HippoRAG and A-Mem have the two largest mean stores but sharply different performance. BM25 has effectively zero LLM and embedding ingestion cost and achieves 43.22% aggregate accuracy; only Hindsight surpasses it, at substantially higher ingestion cost.

  10. Knowl 10 — The benchmark is limited to English, text-only workplace conversations

    limitation

    GroupMemBench covers English-language, text-only workplace conversations in Technology, Healthcare, Manufacturing, and Finance. It does not evaluate multilingual interactions, where speaker-conditioned vocabulary may interact with code-switching or translation drift, or multimodal group conversations containing images, attachments, or voice notes. These settings are identified as future extensions of the graph-grounded synthesis approach, not as capabilities demonstrated by the current benchmark.

Coverage note — The per-domain accuracy breakdowns and the detailed Finance worked example are omitted as separate knowls because they elaborate the aggregate results and failure mechanisms already captured here; related-work comparisons and broader-impact discussion are outside the benchmark's core methodological and empirical contributions.

References

  1. 1.Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, and Ling Yang. Openclaw-rl: Train any agent simply by talking. arXiv preprint arXiv:2603.10165, 2026.
  2. 2.Renjun Xu and Yang Yan. Agent skills for large language models: Architecture, acquisition, security, and the path forward. arXiv preprint arXiv:2602.12430, 2026.
  3. 3.Sheryl Wei Ting Ng and Renwen Zhang. Trust in ai chatbots: A systematic review. Telematics and Informatics, 97:102240, 2025.
  4. 4.Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al. Personal llm agents: Insights and survey about the capability, efficiency and security. arXiv preprint arXiv:2401.05459, 2024.
  5. 5.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345, 2024.
  6. 6.Jess Stratton. An introduction to microsoft copilot. In Copilot for Microsoft 365: harness the power of generative AI in the Microsoft apps you use every day, pages 19–35. Springer, 2024.
  7. 7.Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang, Yuanchen Bei, Yankai Chen, Tao Feng, Xinyu Pan, Zhen Tan, Yu Wang, Tianxin Wei, et al. Rethinking memory mechanisms of foundation agents in the second half. arXiv preprint arXiv:2602.06052, 2026.
  8. 8.Chris Latimer, Nicoló Boschi, Andrew Neeser, Chris Bartholomew, Gaurav Srivastava, Xuan Wang, and Naren Ramakrishnan. Hindsight is 20/20: Building agent memory that retains, recalls, and reflects. arXiv preprint arXiv:2512.12818, 2025.
  9. 9.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025.
  10. 10.Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, et al. Memoryarena: Benchmarking agent memory in interdependent multi-session agentic tasks. arXiv preprint arXiv:2602.16313, 2026.
  11. 11.Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan, Zhongming Yu, Haozhou Xu, Lanxiang Hu, Abhilash Shankarampeta, Zimeng Huang, Wentao Ni, et al. Ama-bench: Evaluating long-horizon memory for agentic applications. arXiv preprint arXiv:2602.22769, 2026.
  12. 12.Qingyao Ai, Yichen Tang, Changyue Wang, Jianming Long, Weihang Su, and Yiqun Liu. Memorybench: A benchmark for memory and continual learning in llm systems. arXiv preprint arXiv:2510.17281, 2025.
  13. 13.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. Long-memeval: Benchmarking chat assistants on long-term interactive memory. arXiv preprint arXiv:2410.10813, 2024.
  14. 14.Chris Frith and Uta Frith. Theory of mind. Current biology, 15(17):R644–R645, 2005.
  15. 15.Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. Fantom: A benchmark for stress-testing machine theory of mind in interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14397–14413, 2023.
  16. 16.Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, and Kuniko Saito. Tomato: Verbalizing the mental states of role-playing llms for benchmarking theory of mind. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 1520–1528, 2025.
  17. 17.Herbert H Clark and Susan E Brennan. Grounding in communication. 1991.
  18. 18.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474, 2020.
  19. 19.Bernal J Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. Advances in neural information processing systems, 37:59532–59569, 2024.
  20. 20.Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. Evaluating very long-term conversational memory of llm agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13851–13870, 2024.
  21. 21.Yuanzhe Hu, Yu Wang, and Julian McAuley. Evaluating memory in llm agents via incremental multi-turn interactions. arXiv preprint arXiv:2507.05257, 2025.
  22. 22.Chuanrui Hu, Tong Li, Xingze Gao, Hongda Chen, Dannong Xu, Yi Bai, Tianwei Lin, Xinda Zhao, Xiaohong Li, Jiaqi An, et al. Evermembench: Benchmarking long-term interactive memory in large language modelsevermembench: Benchmarking long-term interactive memory in large language models. arXiv preprint arXiv:2602.01313, 2026.
  23. 23.Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pages 19724–19731, 2024.
  24. 24.Zeyu Zhang, Quanyu Dai, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. A survey on the memory mechanism of large language model-based agents. ACM Transactions on Information Systems, 43(6):1–47, 2025.
  25. 25.Zhen Tan, Jun Yan, I-Hung Hsu, Rujun Han, Zifeng Wang, Long Le, Yiwen Song, Yanfei Chen, Hamid Palangi, George Lee, et al. In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8416–8439, 2025.
  26. 26.Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, et al. Personamem-v2: Towards personalized intelligence via learning implicit user personas and agentic memory. arXiv preprint arXiv:2512.06688, 2025.
  27. 27.Dongming Jiang, Yi Li, Guanpeng Li, and Bingzhe Li. Magma: A multi-graph based agentic memory architecture for ai agents. arXiv preprint arXiv:2601.03236, 2026.
  28. 28.Charles Packer, Vivian Fang, Shishir_G Patil, Kevin Lin, Sarah Wooders, and Joseph_E Gonzalez. Memgpt: towards llms as operating systems. 2023.
  29. 29.Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110, 2025.
  30. 30.Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, and Zhenhua Dong. Membench: Towards more comprehensive evaluation on the memory of llm-based agents. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19336–19352, 2025.
  31. 31.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing, 2023.
  32. 32.Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024.

Citation

MLA
Yang, J., et al. “GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations”. arXiv, 2026, http://arxiv.org/abs/2605.14498v2.
APA
Yang, J., Lai, K.-H., Wang, X., Chang, S., Harari, Y., & Gabrilovich, E. (2026). GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations. arXiv. http://arxiv.org/abs/2605.14498v2
Chicago
Yang, J., K.-H. Lai, X. Wang, S. Chang, Y. Harari, and E. Gabrilovich. 2026. “GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations”. arXiv. http://arxiv.org/abs/2605.14498v2.
Harvard
Yang, J. et al. (2026) “GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.14498v2.
Vancouver
1. Yang J, Lai K-H, Wang X, Chang S, Harari Y, Gabrilovich E (2026) GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations. arXiv

BibTeX

@article{yang2026groupmembench,
  title = {GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations},
  author = {Yang, Jingbo and Lai, Kwei-Herng and Wang, Xiaowen and Chang, Shiyu and Harari, Yaar and Gabrilovich, Evgeniy},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.14498v2},
  eprint = {2605.14498}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/