Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-party conversations

A multi-party conversation is an interactive dialogue involving three or more participants, such as human speakers, computational agents, or a mix of both. Unlike dyadic interactions between two individuals, multi-party conversations feature complex communicative dynamics characterized by non-linear turn-taking, shifting addressees, multiple concurrent discussion threads, and distinct speaker roles and perspectives. Processing and modeling these exchanges requires identifying who is speaking to whom, maintaining separate context and memory for each participant, and interpreting multimodal signals, including spoken language, vocal acoustics, and visual expressions across all interlocutors.

3 items

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Shiyu Chang, Yaar Harari, Evgeniy Gabrilovich

OrganizationsMicrosoftUniversity of California, Santa Barbara

Why you should read this

Introduces GroupMemBench, a benchmark for evaluating LLM agent memory in multi-party conversations, revealing that current memory systems struggle with speaker-grounded context and collapse to an average accuracy of only 46%.

Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems that extract, retrieve, and apply information across long-running conversations. However, both existing memory systems and benchmarks are built around the dyadic, single-user setup, even though real deployments routinely span groups and channels with multiple users interacting with the agent and with each other. This mismatch leaves three properties of group memory unmeasured: (i) group dynamics that go beyond concatenated one-on-one chats, (ii) speaker-grounded belief tracking, where the per-user memory modeling is needed, and (iii) audience-adapted language, where Theory-of-Mind shifts produce role-specific vocabulary. We introduce GroupMemBench, a benchmark that exposes all three. A graph-grounded synthesis pipeline produces multi-party conversations with controllable reply structure and conditions each message on per-user personas and target audiences. An adversarial query pipeline then binds every question to a specific asker across six categories, spanning multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention, and iteratively searches challenging, realistic queries that reflect comprehensive memory capability. Benchmarking leading memory systems exposes a sharp collapse: the strongest one reaches only 46.0% average accuracy, with knowledge update at 27.1% and term ambiguity at 37.7%, while a simple BM25 baseline matches or exceeds most agent memory systems. This indicates current memory ingestion erases the structural and lexical features group memory depends on, leaving multi-user memory far from solved.

Added

2026-09-30

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

Wenjie Zheng, Jianfei Yu, Rui Xia, Shijin Wang

OrganizationsiFLYTEKNanjing University of Science and TechnologyState Key Laboratory of Cognitive Intelligence

Why you should read this

Proposes a two-stage multimodal framework that isolates the true speaker's face sequence from complex multi-party video scenes to accurately guide conversational emotion recognition via multi-task learning.

Multimodal Emotion Recognition in Multi-party Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information. Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC. However, given an utterance, the face sequence extracted by previous methods may contain multiple people’s faces, which will inevitably introduce noise to the emotion prediction of the real speaker. To tackle this issue, we propose a two-stage framework named Facial expression-aware Multimodal Multi-Task learning (FacialMMT). Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching. With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning. Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. The source code is publicly released at https://github.com/NUSTM/FacialMMT.

Added

2026-09-26