LM Agents for Coordinating Multi-User Information Gathering
Harsh JhamtaniJacob AndreasBen Van Durme
Introduces PeopleJoin, a benchmark for evaluating language model agents on their ability to identify relevant teammates, gather distributed information across multi-user organizations, and synthesize solutions for collaborative tasks.
In modern organizations, critical information is routinely fragmented across multiple employees, requiring significant time and coordination to locate the right people, ask precise questions, and synthesize answers. While large language model agents are increasingly used to assist individual knowledge workers, extending them to coordinate multi-user workplace collaboration introduces major challenges in dealing with incomplete organizational knowledge, complex planning, and human communication overhead.
The article introduces and evaluates PEOPLEJOIN, a novel benchmark designed to assess how effectively language model agents can mediate collaborative information gathering across simulated organizations of 2 to 20 users.
The evaluation approach repurposes standard question-answering and multi-document summarization benchmarks to create synthetic enterprise environments. The authors tested reactive agent architectures using several state-of-the-art language models across 500 question-answering tasks and 200 document-creation tasks. The agents were evaluated on their accuracy in answering queries, their precision and recall in contacting the correct colleagues, and their communication efficiency measured by message counts and word volume. The benchmark was complemented by a 100-task case study involving real human collaborators to validate the simulation environment.
The findings show that current language model agents struggle substantially with multi-user collaboration. Across question-answering tasks, the best-performing model achieved an answer match score of only 54.8 out of 100. Key failure points included poor question formulation that caused colleagues to withhold information (25% of errors), failing to identify or contact all necessary teammates (50% of errors), and navigating organizational redirections, where the match score fell to 38.0. Incorporating a reflection step improved question-answering performance and source precision, though it provided minimal benefit in document creation. In the human-in-the-loop study, human participants asked more clarification questions, leading to slightly higher message counts and marginally better accuracy compared to the fully simulated setup.
These results indicate that deploying fully autonomous agents in collaborative workplace settings presents real operational and productivity risks. Beyond potential privacy concerns when accessing distributed documents, agents that ask vague or excessive questions risk overwhelming colleagues and causing communication fatigue. To manage these trade-offs, organizations considering collaborative agents should adopt human-in-the-loop controls, such as requiring user approval before initiating outreach, and invest in agent mechanisms that learn organizational structures over time to minimize unnecessary messaging.
The benchmark has limitations: it currently focuses solely on English, models only one-on-one text exchanges rather than group channels, and does not fully account for real-world workplace dynamics such as varying colleague response times, availability, and social hierarchies. Readers should view the findings as a credible baseline demonstrating that substantial research into query planning and organizational learning is needed before enterprise-wide autonomous deployment is viable.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). It establishes foundational multi-agent conversational patterns and role-based interaction architectures that inform LM agent coordination in multi-user settings.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It introduces cooperative communicative agent role-playing frameworks essential for understanding conversational information exchange among autonomous agents.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). It investigates Theory of Mind and perspective-taking in multi-agent LLM systems, which underpins identifying teammates and inferring hidden distributed knowledge.
- Paper: Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication, Zhangyue Yin et al. (2023). It analyzes multi-agent communication topologies and cross-model reasoning exchange that motivate structured coordination protocols across synthetic organizations.
- Paper: ProAgent: Building Proactive Cooperative Agents with Large Language Models, Ceyao Zhang et al. (2024). It formalizes proactive intention inference and dynamic coordination mechanisms in multi-agent collaboration without joint prior training.
- Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task, Tao Yu et al. (2018). It provides the underlying cross-domain tabular question-answering formulation from which the PeopleJoin-QA evaluation domain is adapted.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). It defines multi-turn conversational question answering and context-tracking dynamics central to multi-user information-gathering dialogues.
- Paper: Decentralized Multi-Agent Systems with Shared Context, Yuzhen Mao et al. (2026). It extends decentralized multi-agent architectures by introducing shared, verified context mechanisms to eliminate coordination bottlenecks in complex multi-document reasoning.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). It builds upon multi-agent coordination by training a dedicated meta-agent to orchestrate task division and communication among heterogeneous worker models.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). It systematically investigates the scaling principles, communication topologies, and overhead trade-offs of multi-agent team coordination evaluated in collaborative benchmarks.
- Paper: YES AND: A Generative AI Multi-Agent Framework for Enhancing Diversity of Thought in Individual Ideation for Problem-Solving Through Confidence-Based Agent Turn-Taking, Pratik Ghosh et al. (2025). It applies conversational multi-agent problem-solving to organizational ideation using confidence-based turn-taking protocols.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It advances agent evaluation methodologies by deploying autonomous agentic judges to evaluate multi-step collaborative and software workflows across full execution trajectories.
