LM Agents for Coordinating Multi-User Information Gathering

Harsh JhamtaniJacob AndreasBen Van Durme

article2025arXiv7 citations

Introduces PeopleJoin, a benchmark for evaluating language model agents on their ability to identify relevant teammates, gather distributed information across multi-user organizations, and synthesize solutions for collaborative tasks.

Listen

In modern organizations, critical information is routinely fragmented across multiple employees, requiring significant time and coordination to locate the right people, ask precise questions, and synthesize answers. While large language model agents are increasingly used to assist individual knowledge workers, extending them to coordinate multi-user workplace collaboration introduces major challenges in dealing with incomplete organizational knowledge, complex planning, and human communication overhead.

The article introduces and evaluates PEOPLEJOIN, a novel benchmark designed to assess how effectively language model agents can mediate collaborative information gathering across simulated organizations of 2 to 20 users.

The evaluation approach repurposes standard question-answering and multi-document summarization benchmarks to create synthetic enterprise environments. The authors tested reactive agent architectures using several state-of-the-art language models across 500 question-answering tasks and 200 document-creation tasks. The agents were evaluated on their accuracy in answering queries, their precision and recall in contacting the correct colleagues, and their communication efficiency measured by message counts and word volume. The benchmark was complemented by a 100-task case study involving real human collaborators to validate the simulation environment.

The findings show that current language model agents struggle substantially with multi-user collaboration. Across question-answering tasks, the best-performing model achieved an answer match score of only 54.8 out of 100. Key failure points included poor question formulation that caused colleagues to withhold information (25% of errors), failing to identify or contact all necessary teammates (50% of errors), and navigating organizational redirections, where the match score fell to 38.0. Incorporating a reflection step improved question-answering performance and source precision, though it provided minimal benefit in document creation. In the human-in-the-loop study, human participants asked more clarification questions, leading to slightly higher message counts and marginally better accuracy compared to the fully simulated setup.

These results indicate that deploying fully autonomous agents in collaborative workplace settings presents real operational and productivity risks. Beyond potential privacy concerns when accessing distributed documents, agents that ask vague or excessive questions risk overwhelming colleagues and causing communication fatigue. To manage these trade-offs, organizations considering collaborative agents should adopt human-in-the-loop controls, such as requiring user approval before initiating outreach, and invest in agent mechanisms that learn organizational structures over time to minimize unnecessary messaging.

The benchmark has limitations: it currently focuses solely on English, models only one-on-one text exchanges rather than group channels, and does not fully account for real-world workplace dynamics such as varying colleague response times, availability, and social hierarchies. Readers should view the findings as a credible baseline demonstrating that substantial research into query planning and organizational learning is needed before enterprise-wide autonomous deployment is viable.

Cover for LM Agents for Coordinating Multi-User Information Gathering

Abstract

This paper introduces PeopleJoin, a benchmark for evaluating LM-mediated collaborative problem solving. Given a user request, PeopleJoin agents must identify teammates who might be able to assist, converse with these teammates to gather information, and finally compile a useful answer or summary for the original user. PeopleJoin comprises two evaluation domains: PeopleJoin-QA, focused on questions about tabular data, and PeopleJoin-DocCreation, focused on document creation tasks. The two domains are adapted from existing NLP benchmarks for database question answering and multi-document summarization; here, however, the information needed to complete these tasks is distributed across synthetic ``organizations'' of 2--20 users, simulating natural multi-user collaboration scenarios. We implemented several popular LM agent architectures, evaluating their accuracy and efficiency at completing tasks, and highlight new research questions that can be studied using PeopleJoin.

Table of Contents

  • 1 Introduction
  • 2 Challenges in Effectively Steering Multi-User Information Gathering
  • 3 Data
  • 3.1 PeopleJoin-QA
  • 3.2 PeopleJoin-DocCreation
  • 4 Baseline Agent Architectures
  • 4.1 Actions
  • 4.2 Observations and Reflection
  • 4.3 Prompt Structure
  • 5 Evaluation
  • 5.1 Outcome Metrics
  • 5.2 Efficiency Metrics
  • 5.3 Information Source Metrics
  • 6 Experiments
  • 6.1 Results on PeopleJoin-QA
  • 6.2 Results on PeopleJoin-DocCreation
  • 6.3 Case Study with Human Participants
  • 7 Related Work
  • 8 Conclusions and Future Directions
  • References
  • A Additional details on approach
  • A.1 Action descriptions
  • A.2 Exemplars
  • A.3 Action parsing failures
  • A.4 Overview of the prompt structure
  • B Additional details on Experiment Setup
  • B.1 Match score
  • B.2 User Simulators
  • B.3 Qualitative Examples
  • B.4 Human Evaluation Study
  • C Additional details on datasets

Knowls

  1. Knowl 1 — PEOPLEJOIN Benchmark for Multi-User Collaborative Information Gathering

    definition

    PEOPLEJOIN is an evaluation framework designed to assess the capability of large language model (LM) agents to assist users with collaborative, multi-user information gathering and task execution. In a PEOPLEJOIN task, an agent operates within a synthetic organization composed of 2 to 20 users. The problem setup features:

    1. Initiating User: A user who issues a complex request to their digital agent.
    2. Information Fragmentation and Privilege: Essential documents required to fulfill the request are distributed privately across various collaborators within the organization, such that the initiating user's documents alone are insufficient.
    3. Partial Observability: The agent has direct access only to the initiating user's local documents. To identify other knowledgeable collaborators, the agent queries an organization-wide people search interface containing potentially imprecise or high-level descriptions of collaborator expertise.
    4. Communication and Coordination: The agent interacts with other users via a messaging interface, asking targeted questions and managing multi-turn exchanges to collect, join, or summarize the needed information while minimizing communication overhead.

    The benchmark comprises two task families: PEOPLEJOIN-QA (relational question answering across distributed tabular data) and PEOPLEJOIN-DOCCREATION (multi-document aggregation and open-ended text summarization).

  2. Knowl 2 — PEOPLEJOIN-QA Dataset Construction and Structural Challenges

    experimental setup

    PEOPLEJOIN-QA adapts the SPIDER text-to-SQL benchmark into a multi-user collaborative question-answering task across 200 databases. Each SPIDER database represents an organization with 2 to 20 users, where database tables are converted into documents (formatted as sequences of JSON objects per row) and distributed among users such that each user holds a single document and no two users hold the same document.

    To simulate realistic collaborative challenges, three structural transformations are introduced:

    1. Split Documents: A randomly selected table is partitioned equally across two distinct users, requiring the agent to query and aggregate data across multiple individuals (present in 25% of test instances).
    2. Redirection: Scenarios where a redirecting user lacks direct access to the requested table but holds an organizational directory document indicating which target user possesses it, forcing the agent to navigate knowledge hierarchies (present in 22% of test instances).
    3. Missing Information: A randomly chosen table is completely omitted from the organization, rendering the query unanswerable and testing the agent's ability to recognize information absence (present in 9% of test instances).

    Collaborator search hints are generated by taking table and column names and transforming them via few-shot prompting with GPT-4 into simplified, occasionally underspecified natural language descriptions. The benchmark contains 500 test tasks where completing a query requires contacting an average of 1.54 collaborators (variance of 1.12, range 0 to 5).

  3. Knowl 3 — PEOPLEJOIN-DOCCREATION Dataset Construction

    experimental setup

    PEOPLEJOIN-DOCCREATION adapts the MULTINEWS multi-document summarization benchmark to evaluate open-ended document gathering and synthesis across multiple users. Multiple MULTINEWS instances are combined into synthetic organizations of 1 to 7 users (average 5.1 users, variance 4.5), with an average of 6.4 total documents per organization (1.25 documents per user).

    Each organization possesses news articles across 3 topics randomly distributed among its users, meaning some users hold articles on a single topic, some hold articles on multiple topics, and others possess no relevant documents. Collaborator expertise profiles in the people search interface are created by querying GPT-4 on user documents to extract topical keyword lists (e.g., governor election, GOP, health care).

    The evaluation set consists of 200 test instances across 67 organizations. A task requires an agent to identify relevant individuals, query them to retrieve source articles or relevant excerpts, and compile them into a target summary that is derived from an average of 2.7 documents (variance 1.1).

  4. Knowl 4 — Event-Driven Reactive Agent Architecture for Multi-User Collaboration

    model/method

    The reference agent framework uses an event-driven ReAct loop where incoming messages from users trigger iterative loops of action prediction, observation processing, and internal reflection until the agent pauses or concludes the session.

    The action space comprises the following programmatic functions:

    • EnterpriseSearch.search_documents(query: str) -> tuple[str, ...]: BM25 retrieval over documents accessible to the initiating user, returning up to 3 highest-scoring documents.
    • EnterpriseSearch.search_relevant_people(query: str) -> str: BM25 retrieval over employee profile descriptions and knowledge areas, returning up to 10 matching candidate profiles.
    • Enterprise.resolve_primary_user() -> str: Returns metadata (user ID, full name, email) of the initiating user.
    • Enterprise.resolve_person(name: str) -> str: Matches a name string to an organization member and returns their user ID and contact details.
    • Enterprise.send_message(user_id: str, content: str, title: str | None) -> None: Dispatches a textual message to a specified user.
    • Enterprise.send_session_completed() -> None: Closes the task session once the initiating user's goal is resolved.
    • System.finish() -> None: Marks the current turn as complete to await user responses.
    • Reflection.thought(thought: str) -> None: A scratchpad action containing chain-of-thought planning that returns no external observation but remains in the prompt history at future turns (disabled in Reactive-NoRef variants).

    Prompts contain functional action specifications, four few-shot exemplar dialogues exhibiting phenomena like redirection and unanswerable requests, and the full incremental history of events, actions, and observations. If action parsing fails, a formatting reminder is appended up to three times before termination.

  5. Knowl 5 — Evaluation Metrics Suite for Multi-User Collaborative Agents

    definition

    PEOPLEJOIN evaluates agent performance across three dimensions:

    1. Outcome Quality:
      • Match Score (for QA): An automated LLM evaluator (gpt-4-turbo) grades the agent's final response against the gold answer using a discrete 3-point scale: 100100 for a complete match, 5050 for a partial match, and 00 for an incorrect or failed answer. The metric achieves a Cohen's κ=0.81\kappa = 0.81 with human expert annotations.
      • ROUGE-L (for Document Creation): The longest common subsequence ROUGE score comparing the final extracted summary enclosed within delimiter brackets against the ground-truth reference summary.
      • G-Eval (for Document Creation): LM-based evaluations of Relevance, Consistency, and Coherence on a 1–5 scale, provided with access to source documents.
    2. Task Efficiency:
      • Message Count (MsgCnt): The total count of messages exchanged across all participants during the episode.
      • Message Size (MsgSize): The aggregate number of words exchanged across all messages, tokenized using NLTK.
      • People Contacted (#People): The average number of distinct users contacted (including the initiating user).
    3. Information Source Selection:
      • Precision (P-Prec) and Recall (P-Rec): The precision and recall of the set of distinct people contacted by the agent relative to the minimal ground-truth set of people holding the required documents.
  6. Knowl 6 — Empirical Performance and Reflection Ablations on PEOPLEJOIN-QA

    data/table

    Evaluation of baseline reactive agent architectures on the 500 test instances of PEOPLEJOIN-QA across three language model backbones (gpt-4-turbo, gpt-4o, and phi-3-medium), comparing full Reactive prompting against Reactive-NoRef (which disables internal reflection actions):

    Method Match ↑\uparrow MsgCnt ↓\downarrow MsgSize ↓\downarrow #People ↓\downarrow P-Prec ↑\uparrow P-Rec ↑\uparrow
    LLM: gpt-4-turbo
    Reactive 54.8 9.0 193 1.5 0.61 0.89
    Reactive-NoRef 48.0 9.2 187 1.5 0.55 0.82
    LLM: gpt-4o
    Reactive 48.7 9.7 179 1.2 0.60 0.83
    Reactive-NoRef 40.4 10.4 209 2.0 0.52 0.78
    LLM: phi-3-medium
    Reactive 24.4 6.7 122 1.0 0.23 0.52
    Reactive-NoRef 20.0 16.3 295 1.7 0.39 0.62

    The results highlight that multi-user information gathering is difficult even for frontier models, with gpt-4-turbo reaching a peak Match score of only 54.8. gpt-4-turbo consistently outperforms gpt-4o and phi-3-medium. Incorporating reflection (Reactive) consistently improves Match scores and collaborator targeting precision/recall over Reactive-NoRef across model sizes.

  7. Knowl 7 — Empirical Performance on PEOPLEJOIN-DOCCREATION

    data/table

    Performance of agent architectures across language models on the 200 test instances of PEOPLEJOIN-DOCCREATION. G-Eval reports scores for Relevance / Consistency / Coherence:

    Method Rouge ↑\uparrow G-Eval ↑\uparrow MsgCnt ↓\downarrow MsgSize ↓\downarrow P-Prec ↑\uparrow P-Rec ↑\uparrow
    LLM: gpt-4-turbo
    Reactive 16.3 4.00 / 4.16 / 4.07 12.6 1330 0.99 0.88
    Reactive-NoRef 16.5 4.20 / 4.33 / 4.14 12.4 1281 0.97 0.87
    LLM: gpt-4o
    Reactive 12.2 2.99 / 3.33 / 3.00 9.9 1180 0.95 0.80
    Reactive-NoRef 12.6 3.15 / 3.42 / 2.65 10.9 1268 0.90 0.90
    LLM: phi-3-medium
    Reactive 11.5 2.84 / 3.31 / 2.81 11.0 996 0.66 0.69
    Reactive-NoRef 11.3 2.71 / 2.64 / 3.20 11.3 948 0.65 0.67

    For reference, an IdealAgent on this task achieves G-Eval scores of 5.0, MsgCnt of 6.3, MsgSize of 1592, and contacts 1.7 people. Unlike the structured QA setting, reflection (Reactive vs Reactive-NoRef) yields no notable performance advantage on open-ended document creation.

  8. Knowl 8 — Sub-Category Performance Breakdown and Failure Modes in PEOPLEJOIN-QA

    empirical result

    An analysis of the gpt-4-turbo Reactive agent across specific structural sub-categories of PEOPLEJOIN-QA reveals marked differences in task difficulty:

    • Unanswerable queries: 87.5 Match score.
    • Document Split: 50.0 Match score.
    • Redirection: 38.0 Match score.

    While agents excel at identifying unanswerable questions when information is missing, they struggle substantially with information fragmentation and traversing multi-step knowledge hierarchies.

    A manual qualitative failure analysis of 40 randomly sampled imperfect cases identifies four primary error modes:

    1. Incomplete collaborator coverage (30%): The agent failed to contact all relevant individuals, leading to an incorrect final answer.
    2. Suboptimal query formulation (25%): The agent asked poorly worded or overly specific questions that caused collaborators to report they lacked relevant documents when they in fact possessed them.
    3. Premature abandonment (20%): The agent failed to exhaustively search for or contact all relevant users and incorrectly informed the initiating user that the required information could not be gathered.
    4. Orchestration failures (10%): The agent made execution errors, such as failing to invoke search tools for people or documents.
  9. Knowl 9 — Evaluation of Communication Strategies Against Static Baselines

    data/table

    Comparison on PEOPLEJOIN-QA between the Reactive agent and fixed communication baseline strategies using gpt-4-turbo:

    Method Match ↑\uparrow MsgCnt ↓\downarrow P-Prec ↑\uparrow
    Reactive 54.8 9.0 0.61
    MessageAllOnce 34.6 11.4 0.37
    MessageNone 19.2 4.1 N/A
    IdealAgent 100.0 7.0 1.00
    • MessageNone acts as a lower bound by attempting to answer queries using solely the initiating user's local documents without contacting anyone (19.2 Match).
    • MessageAllOnce broadcasts the raw user query to every employee in the organization exactly once (34.6 Match, 11.4 messages).
    • IdealAgent represents theoretical optimal execution (contacting only the true minimal set of users with perfect queries, requiring 7.0 messages and 1.5 people on average).

    The results demonstrate that targeted collaborator selection and iterative, multi-turn dialogue are critical, as indiscriminate broadcasting incurs higher communication costs while yielding inferior accuracy.

  10. Knowl 10 — Human Evaluation Case Study Validating LLM-Based User Simulators

    data/table

    To evaluate the realism and fidelity of gpt-4-turbo-powered user simulators, a human evaluation study was conducted across 100 randomly sampled instances of PEOPLEJOIN-QA. In each test task, one simulated collaborator from the gold set of required people was replaced by a recruited human participant given access to the identical persona and documents:

    Setting Match ↑\uparrow MsgCnt ↓\downarrow MsgSize ↓\downarrow
    Human Participant 50 10.0 198
    Simulation 44 9.3 187

    Interactions with human participants yielded slightly longer responses and more clarification requests, producing a slight increase in total message count (10.0 vs 9.3) and words exchanged (198 vs 187). The resulting Match score was slightly higher with humans (50 vs 44). Overall, the dialogue trajectories and outcomes between simulated users and human collaborators were qualitatively and quantitatively consistent, validating the simulator setup.

Coverage note — Deliberately omitted verbatim prompt templates and full multi-turn execution transcripts from the appendix, as they provide implementation and reproducibility details rather than core conceptual contributions.

References

  1. 1.Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219.
  2. 2.Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O'Reilly Media, Inc.".
  3. 3.Jenna Butler, Sonia Jaffe, Nancy Baym, Mary Czerwinski, Shamsi Iqbal, Kate Nowak, Sean Rintel, Abigail Sellen, Mihaela Vorvoreanu, Najeeb G. Abdulhamid, Judith Amores, Reid Andersen, Kagonya Awori, Maxamed Axmed, danah boyd, James Brand, Georg Buscher, Dean Carignan, Martin Chan, Adam Coleman, Scott Counts, Madeleine Daepp, Adam Fourney, Daniel G. Goldstein, Andy Gordon, Aaron L Halfaker, Javier Hernandez, Jake Hofman, Jenny Lay-Flurrie, Vera Liao, Siân Lindley, Sathish Manivannan, Charlton Mcilwain, Subigya Nepal, Jennifer Neville, Stephanie Nyairo, Jacki O'Neill, Victor Poznanski, Gonzalo Ramos, Nagu Rangan, Lacey Rosedale, David Rothschild, Tara Safavi, Advait Sarkar, Ava Scott, Chirag Shah, Neha Parikh Shah, Teny Shapiro, Ryland Shaw, Auste Simkute, Jina Suh, Siddharth Suri, Ioana Tanase, Lev Tankelevitch, Adam Troy, Mengting Wan, Ryen W. White, Longqi Yang, Brent Hecht, and Jaime Teevan. 2023. Microsoft new future of work report 2023. Technical Report MSR-TR-2023-34, Microsoft.
  4. 4.Alexander Richard Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1074–1084.
  5. 5.Ian Gemp, Yoram Bachrach, Marc Lanctot, Roma Patel, Vibhavari Dasagi, Luke Marris, Georgios Piliouras, and Karl Tuyls. 2024. States as strings as strategies: Steering language models with game-theoretic solvers. arXiv preprint arXiv:2402.01704.
  6. 6.Harsh Jhamtani, Hao Fang, Patrick Xia, Eran Levy, Jacob Andreas, and Ben Van Durme. 2024. Natural language decomposition and interpretation of complex utterances. IJCAI.
  7. 7.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  8. 8.Yujin Kim and Chin-Chia Hsu. 2024. Leveraging large language models for hybrid workplace decision support. Preprint, arXiv:2402.03616.
  9. 9.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero-shot relation extraction via reading comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 333–342, Vancouver, Canada. Association for Computational Linguistics.
  10. 10.Michael Lewis. 1998. Designing for human-agent interaction. Ai magazine, 19(2):67–67.
  11. 11.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  12. 12.Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner. 2024. Decision-oriented dialogue for human-ai collaboration. Transactions of the Association for Computational Linguistics, 12:892–911.
  13. 13.Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating wikipedia by summarizing long sequences. In International Conference on Learning Representations.
  14. 14.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 2511–2522. Association for Computational Linguistics.
  15. 15.Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023. Can llms keep a secret? testing privacy implications of language models via contextual integrity theory. arXiv preprint arXiv:2310.17884.
  16. 16.OpenAI. 2023. GPT-4 technical report. Computing Research Repository, arXiv:2303.08774.
  17. 17.Marios Papachristou, Longqi Yang, and Chin-Chia Hsu. 2023. Leveraging large language models for collective decision-making. arXiv preprint arXiv:2311.04928.
  18. 18.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789.
  19. 19.Amit P Sheth and James A Larson. 1990. Federated database systems for managing distributed, heterogeneous, and autonomous databases. ACM Computing Surveys (CSUR), 22(3):183–236.
  20. 20.Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, and Saab Mansour. 2024. Finesure: Fine-grained summarization evaluation using llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 906–922. Association for Computational Linguistics.
  21. 21.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  22. 22.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. Computing Research Repository, arXiv:2307.09288.
  23. 23.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345.
  24. 24.Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018. Constructing datasets for multi-hop reading comprehension across documents. Transactions of the Association for Computational Linguistics, 6:287–302.
  25. 25.Tomer Wolfson, Mor Geva, Ankit Gupta, Yoav Goldberg, Matt Gardner, Daniel Deutch, and Jonathan Berant. 2020. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8:183–198.
  26. 26.Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. arXiv preprint arXiv:2308.08155.
  27. 27.Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864.
  28. 28.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  29. 29.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  30. 30.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921.
  31. 31.Carlo Zaniolo. 1997. Advanced database systems. Morgan Kaufmann.

Citation

MLA
Jhamtani, H., et al. “LM Agents for Coordinating Multi-User Information Gathering”. arXiv, 2025, http://arxiv.org/abs/2502.12328v1.
APA
Jhamtani, H., Andreas, J., & Durme, B. V. (2025). LM Agents for Coordinating Multi-User Information Gathering. arXiv. http://arxiv.org/abs/2502.12328v1
Chicago
Jhamtani, H., J. Andreas, and B. V. Durme. 2025. “LM Agents for Coordinating Multi-User Information Gathering”. arXiv. http://arxiv.org/abs/2502.12328v1.
Harvard
Jhamtani, H., Andreas, J. and Durme, B.V. (2025) “LM Agents for Coordinating Multi-User Information Gathering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.12328v1.
Vancouver
1. Jhamtani H, Andreas J, Durme BV (2025) LM Agents for Coordinating Multi-User Information Gathering. arXiv

BibTeX

@article{jhamtani2025agents,
  title = {LM Agents for Coordinating Multi-User Information Gathering},
  author = {Jhamtani, Harsh and Andreas, Jacob and Durme, Benjamin Van},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.12328v1},
  eprint = {2502.12328}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/