Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Qineng WangZihao WangYing SuHanghang TongYangqiu Song

article2024ACL224 citations

Demonstrates that single language models with standard few-shot prompts match the reasoning performance of complex multi-agent discussion frameworks, revealing that multi-agent systems only provide advantages in zero-shot settings.

Listen

Recent advances in artificial intelligence have popularized multi-agent discussion frameworks, where multiple AI models debate or converse to solve complex reasoning problems. While proponents argue that simulating group interactions significantly improves problem-solving accuracy over single models, these multi-agent pipelines require substantially higher computational resources, token usage, and engineering complexity. This article addresses whether multi-agent discussion is truly necessary to achieve state-of-the-art reasoning performance or if the observed benefits stem primarily from effective prompt design.

The main objective of the article is to systematically evaluate the performance of single-agent setups against multi-agent discussion frameworks across various reasoning tasks and model architectures. It aims to determine the exact conditions under which multi-agent discussion provides a genuine performance advantage over a well-prompted single model.

To evaluate these methods, the authors developed a novel group-discussion framework called Conquer-and-Merge Discussion (CMD), which organizes agents into sub-groups to reduce token overhead, followed by voting and tie-breaking. They conducted comparative experiments across three standard reasoning benchmarks: ECQA (commonsense reasoning), GSM8K (math word problems), and FOLIO-wiki (deductive logic). The tests evaluated single models and multi-agent frameworks—including Debate, MAD, ReConcile, and CMD—across single-model configurations (such as ChatGPT-3.5) and mixed-model setups incorporating Gemini Pro, Bard, and open-source LLaMA models.

The investigation yielded four critical findings. First, a single AI agent equipped with a strong prompt containing a concrete demonstration achieves essentially the same accuracy as the best multi-agent discussion frameworks (for example, scoring approximately 75.6% on average compared to 70.0%–74.5% across discussions using ChatGPT-3.5). Second, multi-agent discussions show a clear advantage only in zero-shot scenarios where no task demonstrations are provided (averaging roughly 70.8% with CMD versus 67.4% for a direct single agent). Third, multi-agent frameworks introduce unique failure modes, notably wrong answer propagation, where an agent abandons a correct initial answer to follow an incorrect consensus, and judge errors during tie-breaking. Fourth, in mixed-model discussions, stronger models like Gemini Pro effectively elevate the reasoning performance of weaker models like Bard and smaller open-source models across discussion rounds.

These findings have immediate practical implications for cost, infrastructure, and deployment strategy. In enterprise settings where task-specific demonstrations and domain guidance can be engineered into prompts, deploying multi-agent discussion frameworks introduces unnecessary computational costs and operational latency without delivering measurable accuracy gains. However, when tasks are open-ended or high-quality few-shot examples cannot be created, multi-agent interactions serve as an effective alternative to improve reasoning.

Organizations should prioritize investing in high-quality prompt engineering and domain-specific demonstrations for single-agent systems as their primary deployment strategy. Multi-agent discussion setups should be reserved selectively for scenarios lacking curated examples or where heterogeneous model architectures are paired to allow stronger models to guide cheaper, weaker models. When implementing multi-agent pipelines, developers must incorporate safeguards against group conformity errors and voting biases.

The conclusions are supported by structured evaluations on established reasoning benchmarks, but certain limitations remain. The empirical analysis tested three primary commercial models and select open-source variants on a limited set of reasoning datasets, representing simplified agent sessions rather than complex systems with external memory or search tools. Confidence is high regarding the evaluated benchmarks, though stakeholders should validate these trade-offs before scaling multi-agent architectures in production environments.

arXiv: 2402.18272HKUST-KnowComp/LLM-discussion

No sufficiently relevant recommendations were found.

Cover for Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Abstract

Recent progress in LLMs discussion suggests that multi-agent discussion improves the reasoning abilities of LLMs. In this work, we reevaluate this claim through systematic experiments, where we propose a novel group discussion framework to enrich the set of discussion mechanisms. Interestingly, our results show that a single-agent LLM with strong prompts can achieve almost the same performance as the best existing discussion approach on a wide range of reasoning tasks and backbone LLMs. We observe that the multi-agent discussion performs better than a single agent only when there is no demonstration in the prompt. Further study reveals the common interaction mechanisms of LLMs during the discussion.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 2.1 What is Multi-Agent Discussion?
  • 2.2 Existing Discussion Frameworks
  • 3 CMD: Conquer-and-Merge Discussion
  • 3.1 Message-Passing Algorithm
  • 3.2 Three Stages of CMD
  • 4 Experimental Setups
  • 4.1 Implementation Details and Metrics
  • 4.2 Downstream Tasks
  • 5 Experiments on Single LLM
  • 5.1 Analysis of FOLIO-wiki Dataset
  • 5.2 Evaluation on All Tasks
  • 5.3 Two Discussion Error Types: A Case Study
  • 5.4 Summary
  • 6 Experiments on Multiple LLMs
  • 6.1 Validate Findings on Multiple LLMs Scenarios
  • 6.2 Enhancing Agents in Weaker LLMs with Support from Stronger LLMs
  • 7 Related Work
  • 7.1 Prompting LLM for Reasoning
  • 7.2 Multi-agent Discussion for Reasoning with LLMs
  • 8 Conclusion
  • 9 Ethical Considerations
  • 10 Limitations
  • 11 Acknowledgement
  • References
  • A Extended Empirical Results
  • A.1 Open-source language model
  • A.2 CMD and other methods to enhance LLMs
  • B Discussion Engineering and Agent Symmetry
  • B.1 Agent symmetry in discussion engineering
  • C CMD: Conquer and Merge Discussion Framework
  • C.1 Motivation
  • C.2 Problem Definition
  • C.3 CMD Stages
  • C.4 Message-Passing Algorithm
  • D An CMD Example
  • D.1 Meta Prompt
  • D.2 Round 1 Answer
  • D.3 Middle System and User Prompts in Round 1
  • D.4 Round 2 Answer
  • D.5 Middle System Prompt at the End of Round 2
  • D.6 Round 3 Answer
  • E CMD Secretary - A Tie Case Solution
  • F Extended Related Work
  • F.1 Large language models L
  • F.2 Prompt decorator p(·; T ,L) for reasoning
  • F.3 Mechanism M for reasoning

Knowls

  1. Knowl 1 — Conquer-and-Merge Discussion (CMD)

    model/method

    CMD is a multi-agent reasoning framework that organizes agents into discussion groups, then combines their answers by voting. In the usual configuration, agents are assigned to groups of three and discuss the same task for a fixed number of rounds. During a round, an agent receives the previous-round answers and explanations from other members of its own group, but receives only the answers—not the explanations—from agents in other groups. This limits the amount of cross-group discussion content while preserving access to detailed local reasoning.

    After discussion, all active agents cast equally weighted votes for their final viewpoints. A majority determines the answer. If the votes tie, a secretary agent can decide using the competing viewpoints and explanations selected from agents holding each position. CMD also allows a representative-based alternative: one representative per group advances to a higher-level discussion, and this process can repeat until the tie is resolved or one agent remains.

  2. Knowl 2 — Discussion performance depends on whether demonstrations are provided

    data/table

    On ECQA, GSM8K, and FOLIO-wiki, the authors compared a single ChatGPT-3.5 agent with four discussion frameworks, with and without task demonstrations. The direct condition has no demonstration; the demo condition includes one. Scores are accuracy percentages. Without demonstrations, six-agent CMD exceeds the single agent on all three tasks and on the average score. With demonstrations, the single agent has the highest average, while CMD remains close; thus discussion is not uniformly superior when prompts include demonstrations.

    Method ECQA Direct ECQA Demo GSM8K Direct GSM8K Demo FOLIO Direct FOLIO Demo Avg. Direct Avg. Demo
    Single Agent 63.00 67.00 69.00 83.00 70.22 76.09 67.41 75.63
    MAD (3 Agents) 55.00 58.00 74.00 78.00 61.25 74.13 63.42 70.04
    Debate (3 Agents) 67.00 65.00 78.00 81.00 70.00 75.65 71.67 73.88
    Debate (6 Agents) 65.00 64.00 74.00 78.00 69.13 74.78 69.38 72.26
    CMD (6 Agents) 64.00 63.00 75.00 83.00 73.26 77.39 70.75 74.46
  3. Knowl 3 — Prompt components affect FOLIO-wiki accuracy

    data/table

    The FOLIO-wiki prompt ablation with ChatGPT-3.5 varies a detailed question description (Q-Desc.), an answer-format description (A-Desc.), and a task-specific demonstration (Demo.). When a component is disabled, the input contains only the question plus any enabled components. Scores are accuracy percentages. The demonstrated condition shown includes both descriptions; compared with the row having both descriptions but no demonstration, adding the demonstration raises accuracy for every listed method. With that demonstration, CMD scores 77.39%, compared with 76.09% for the single agent. Performance also varies with the prompt components and discussion framework, so no one description component yields the same effect in every setting.

    Q-Desc. A-Desc. Demo. MAD (3) Debate (3) Debate (6) Single Agent CMD (6)
    No No No 64.13 70.00 69.13 70.22 73.26
    Yes No No 74.13 75.65 76.30 73.26 74.13
    No Yes No 68.91 71.96 71.74 71.30 73.89
    Yes Yes No 71.96 70.22 70.00 73.91 71.09
    Yes Yes Yes 74.13 75.65 74.78 76.09 77.39
  4. Knowl 4 — Evaluation conditions and benchmarks

    experimental setup

    The study evaluates reasoning accuracy on ECQA, GSM8K, and a curated FOLIO-wiki dataset. Because of resource constraints, the ECQA and GSM8K evaluations use 100 sampled test instances each; the FOLIO-wiki evaluation covers all 460 instances in its curated version, which excludes cases judged flawed. The principal single-LLM experiments use ChatGPT-3.5, implemented as Azure OpenAI gpt-35-turbo (0613). The multi-LLM experiments also use Gemini Pro and Bard, represented by chat-bison-001. CMD uses a dialogue temperature of 0.25, and discussion frameworks are configured for a maximum of three discussion rounds. Accuracy is the evaluation metric.

  5. Knowl 5 — Multi-LLM discussions match a strong single LLM with demonstrations

    data/table

    The multi-LLM evaluation compares individual Bard, Gemini Pro, and ChatGPT-3.5 agents with ReConcile and CMD discussions using all three LLMs. Direct means no demonstration; Demo means a demonstration is supplied. All entries are accuracy percentages. With demonstrations, CMD's 78.66% average is close to the strongest individual model, Gemini Pro at 78.59%; ReConcile averages 78.36%. Without demonstrations, both discussions exceed the best single-agent average: CMD scores 76.93% and ReConcile 76.11%, versus Gemini Pro's 74.38%. These results reproduce the broad pattern that a strong single agent can rival discussion with demonstrations, while discussions can help when demonstrations are absent.

    Category Method / LLM ECQA Direct ECQA Demo GSM8K Direct GSM8K Demo FOLIO Direct FOLIO Demo Avg. Direct Avg. Demo
    Single Agent Bard 66.00 65.00 47.00 54.00 70.00 71.96 61.00 63.65
    Single Agent Gemini Pro 74.00 75.00 75.00 81.00 74.13 79.78 74.38 78.59
    Single Agent ChatGPT-3.5 63.00 67.00 69.00 83.00 70.22 76.09 67.41 75.63
    Discussion ReConcile (Bard, Gemini, ChatGPT) 70.00 71.00 78.00 83.00 80.34 81.09 76.11 78.36
    Group Discussion CMD (Bard, Gemini, ChatGPT) 73.00 72.00 78.00 82.00 79.78 81.96 76.93 78.66
  6. Knowl 6 — Two recurring failure modes in agent discussions

    empirical result

    The authors identify two ways a discussion can produce a wrong answer even when a single agent answers correctly. A judge mistake occurs when an agent tasked with resolving disagreement chooses the incorrect answer, a risk in judge-based decisions and in CMD's secretary decision when votes tie. Wrong-answer propagation occurs when an agent changes an initially correct answer after exposure to other agents' responses and adopts an incorrect position; the paper describes this as the most common discussion error it observed, including cases where most initial answers were correct. These categories are illustrated through a FOLIO-wiki case study, but the paper does not provide a quantitative error-rate estimate.

  7. Knowl 7 — Weaker agents can improve during discussion with stronger models

    data/table

    Round-level CMD results show that weaker LLM agents can gain accuracy while interacting with stronger agents, although this pattern is not uniform for every participant. The table reports accuracy percentages by round for four configurations involving Gemini Pro, Bard or ChatGPT-3.5, and Llama 2 at 7B or 70B. Llama 2 accuracy rises across rounds in all four configurations; ChatGPT-3.5 also rises in the two configurations where it participates. Stronger agents sometimes lose accuracy, especially when paired with Llama 2 7B, so the evidence supports improvement of weaker agents rather than universal improvement for every agent.

    CMD agent combination Agent Round 1 Round 2 Round 3
    Gemini Pro, Bard, Llama 2 7B Gemini Pro 68.00% 58.50% 55.50%
    Bard 47.00% 45.50% 39.00%
    Llama 2 7B 22.00% 30.50% 32.00%
    Gemini Pro, Bard, Llama 2 70B Gemini Pro 65.00% 64.00% 61.00%
    Bard 46.50% 43.00% 40.50%
    Llama 2 70B 41.50% 51.50% 51.50%
    Gemini Pro, Llama 2 7B, ChatGPT-3.5 Gemini Pro 67.00% 66.50% 68.50%
    Llama 2 7B 22.50% 28.50% 36.50%
    ChatGPT-3.5 68.50% 72.50% 73.50%
    Gemini Pro, Llama 2 70B, ChatGPT-3.5 Gemini Pro 66.50% 73.00% 76.00%
    Llama 2 70B 42.50% 67.00% 71.50%
    ChatGPT-3.5 70.00% 75.50% 77.50%
  8. Knowl 8 — Message synchronization for discussion frameworks

    algorithm

    The message-passing procedure, MesSync, is designed to synchronize communication across discussion rules and agents that may use different LLMs or inference methods. Its inputs are a discussion rule RR, agents AA, an agent-attribute table TT that maps each agent to its response-generating method, and initial prompt messages MM. A message records its content, sender, intended receivers, and discussion depth. The output is the completed interaction; the discussion rule determines when it is over and how messages are merged, transformed, validated, and routed.

    Input: Discussion rule R, agents A, bot/method table T, initial messages M
    Output: Completed discussion under R
    Initialize message queue Qmsg with M
    Initialize outgoing queue Qsend as empty
    Set speaker to R's first speaker and depth to 0
    While Qmsg is not empty or R is not over:
        If Qmsg is empty:
            Add a silence message at the current depth
        Set depth to the depth of the first queued message
        Remove all messages at that depth as a batch
        For each agent in A:
            Merge the batch messages intended for that agent using R
            Add the merged message to Qsend
        Initialize a buffer H for messages marked HOLD
        For each outgoing message at the current dispatch depth:
            If the message is marked HOLD:
                Store its content in H under its message name
            Otherwise:
                Set the speaker to the message name
                Combine the message content with any buffered content for that speaker
                Transform the combined input using R
                Generate a response using that speaker's method in T
                Validate the response using R
                Obtain the receivers from R
                If receivers are not empty:
                    Enqueue the response with sender, receivers, and next depth
        If R is over:
            Stop
    Return the completed interaction

    The paper does not specify a complexity bound; routing and validation behavior are delegated to the supplied discussion rule.

  9. Knowl 9 — Six-sample self-consistency is comparable to demonstrations and CMD

    data/table

    On GSM8K, FOLIO-wiki, and ECQA, the authors compare standard chain-of-thought prompting, a demonstrated single-agent prompt, six-trial self-consistency, and six-agent CMD. Scores are accuracy percentages. Self-consistency uses a prompt with demonstrations and achieves 80% on GSM8K, 76.96% on FOLIO-wiki, and 68% on ECQA. Its performance is close to the demonstrated single-agent and CMD results, supporting the paper's claim that strong single-agent prompting or repeated single-agent inference can reach performance similar to discussion in these settings.

    Method GSM8K FOLIO-wiki ECQA
    Standard CoT 69% 70.22% 63%
    Demonstrations 83% 76.09% 67%
    Self-Consistency (6) 80% 76.96% 68%
    CMD (6) 83% 77.39% 63%
  10. Knowl 10 — Scope and stated limitations

    limitation

    The study evaluates reasoning benchmarks and tests only three proprietary LLMs in its main multi-LLM experiments, citing computational and financial constraints; the authors note that broader model coverage is needed to assess generalizability and scalability. They also treat each LLM session as an agent, a simplified agent design that does not incorporate more elaborate reasoning procedures, external tools, or knowledge bases. The authors identify broader applications beyond reasoning tasks—including strategic planning and interactive games—as directions for future evaluation.

Coverage note — The appendix's formal computational-graph and agent-symmetry framework, plus the illustrative CMD transcript, are omitted as ancillary formalization and example material rather than central empirical or procedural contributions.

References

  1. 1.Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. Explanations for commonsenseqa: New dataset and models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3050–3065.
  2. 2.Philip W Anderson. 1972. More is different: Broken symmetry and the nature of the hierarchical structure of science. Science, 177(4047):393–396.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  4. 4.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, et al. 2023. Graph of thoughts: Solving elaborate problems with large language models. arXiv preprint arXiv:2308.09687.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  7. 7.Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. 2023a. Reconcile: Round-table conference improves reasoning via consensus among diverse llms. arXiv preprint arXiv:2309.13007.
  8. 8.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023b. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  11. 11.Constantinos Daskalakis and Seth Matthew Weinberg. 2012. Symmetries and optimal multi-dimensional mechanism design. In Proceedings of the 13th ACM conference on Electronic commerce, pages 370–387.
  12. 12.Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang. 2023. Active prompting with chain-of-thought for large language models. arXiv preprint arXiv:2302.12246.
  13. 13.Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023. Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325.
  14. 14.Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. 2022. Successive prompting for decomposing complex questions. arXiv preprint arXiv:2212.04092.
  15. 15.Weizhi Fei, Xueyan Niu, Pingyi Zhou, Lu Hou, Bo Bai, Lei Deng, and Wei Han. 2023. Extending context window of large language models via semantic compression. arXiv preprint arXiv:2312.09571.
  16. 16.Weizhi Fei, Zihao Wang, Hang Yin, Yang Duan, Hanghang Tong, and Yangqiu Song. 2024. Soft reasoning on uncertain knowledge graphs. arXiv preprint arXiv:2403.01508.
  17. 17.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  18. 18.Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, et al. 2022. Folio: Natural language reasoning with first-order logic. arXiv preprint arXiv:2209.00840.
  19. 19.Shima Imani, Liang Du, and Harsh Shrivastava. 2023. Mathprompter: Mathematical reasoning using large language models. arXiv preprint arXiv:2303.05398.
  20. 20.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406.
  21. 21.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  22. 22.Jean-Jacques Laffont and David Martimort. 2000. Mechanism design with collusion and correlation. Econometrica, 68(2):309–342.
  23. 23.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  24. 24.Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118.
  25. 25.Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su. 2023. Deductive verification of chain-of-thought reasoning. arXiv preprint arXiv:2306.03872.
  26. 26.Lihui Liu, Zihao Wang, Ruizhong Qiu, Yikun Ban, and Hanghang Tong. 2024. Logic query of thoughts: Guiding large language models to answer complex logic queries with knowledge graphs. arXiv preprint arXiv:2404.04264.
  27. 27.Ruibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu, Soroush Vosoughi, Claire Cui, Denny Zhou, and Andrew M Dai. 2022a. Mind’s eye: Grounded language model reasoning through simulation. arXiv preprint arXiv:2210.05359.
  28. 28.Zhixuan Liu, Zihao Wang, Yuan Lin, and Hang Li. 2022b. A neural-symbolic approach to natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 2159–2172.
  29. 29.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842.
  30. 30.Aman Madaan, Niket Tandon, Peter Clark, and Yiming Yang. 2022. Memory-assisted prompt editing to improve gpt-3 after deployment. arXiv preprint arXiv:2201.06009.
  31. 31.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651.
  32. 32.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  33. 33.Marvin Minsky. 1988. Society of mind. Simon and Schuster.
  34. 34.OpenAI. 2022. Chatgpt. https://openai.com/blog/chatgpt.
  35. 35.OpenAI. 2023. Gpt-4 technical report.
  36. 36.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  37. 37.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  38. 38.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems.
  39. 39.Kristopher Tapp. 2021. Symmetry. Springer.
  40. 40.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  41. 41.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  42. 42.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  43. 43.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023a. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091.
  44. 44.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022a. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  45. 45.Zihao Wang, Weizhi Fei, Hang Yin, Yangqiu Song, Ginny Wong, and Simon See. 2023b. Wasserstein-fisher-rao embedding: Logical query embeddings with local comparison and global transport. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13679–13696.
  46. 46.Zihao Wang, Yangqiu Song, Ginny Wong, and Simon See. 2023c. Logical message passing networks with one-hop inference on atomic formulas. In The Eleventh International Conference on Learning Representations.
  47. 47.Zihao Wang, Hang Yin, and Yangqiu Song. 2021. Benchmarking the combinatorial generalizability of complex query answering on knowledge graphs. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  48. 48.Zihao Wang, Hang Yin, and Yangqiu Song. 2022b. Logical queries on knowledge graphs: Emerging interface of incomplete relational data. Data Engineering, page 3.
  49. 49.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  50. 50.Lilian Weng. 2023. Llm-powered autonomous agents. lilianweng.github.io.
  51. 51.Yixuan Weng, Minjun Zhu, Shizhu He, Kang Liu, and Jun Zhao. 2022. Large language models are reasoners with self-verification. arXiv preprint arXiv:2212.09561.
  52. 52.Zhiheng Xi, Senjie Jin, Yuhao Zhou, Rui Zheng, Songyang Gao, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. Self-polish: Enhance reasoning in large language models via problem refinement. arXiv preprint arXiv:2305.14497.
  53. 53.Fangzhi Xu, Qika Lin, Jiawei Han, Tianzhe Zhao, Jun Liu, and Erik Cambria. 2023a. Are large language models really good logical reasoners? a comprehensive evaluation from deductive, inductive and abductive views. arXiv preprint arXiv:2306.09841.
  54. 54.Xiaohan Xu, Chongyang Tao, Tao Shen, Can Xu, Hongbo Xu, Guodong Long, and Jian-guang Lou. 2023b. Re-reading improves reasoning in language models. arXiv preprint arXiv:2309.06275.
  55. 55.Yao Xu, Shizhu He, Jiabei Chen, Zihao Wang, Yangqiu Song, Hanghang Tong, Kang Liu, and Jun Zhao. 2024. Generate-on-graph: Treat llm as both agent and kg in incomplete knowledge graph question answering. arXiv preprint arXiv:2404.14741.
  56. 56.Tianci Xue, Ziqi Wang, Zhenhailong Wang, Chi Han, Pengfei Yu, and Heng Ji. 2023. Rcot: Detecting and rectifying factual inconsistency in reasoning by reversing chain-of-thought. arXiv preprint arXiv:2305.11499.
  57. 57.Zhicheng Yang, Jinghui Qin, Jiaqi Chen, Liang Lin, and Xiaodan Liang. 2022. Logicsolver: Towards interpretable math word problem solving with logical prompt-enhanced learning. arXiv preprint arXiv:2205.08232.
  58. 58.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  59. 59.Yao Yao, Zuchao Li, and Hai Zhao. 2023b. Beyond chain-of-thought, effective graph-of-thought reasoning in large language models. arXiv preprint arXiv:2305.16582.
  60. 60.Hang Yin, Zihao Wang, and Yangqiu Song. 2024a. Meta operator for complex query answering on knowledge graphs. arXiv preprint arXiv:2403.10110.
  61. 61.Hang Yin, Zihao Wang, and Yangqiu Song. 2024b. Rethinking complex queries on knowledge graphs with neural link predictors. In The Twelfth International Conference on Learning Representations.
  62. 62.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488.
  63. 63.Jintian Zhang, Xin Xu, and Shumin Deng. 2023a. Exploring collaboration mechanisms for llm agents: A social psychology view. arXiv preprint arXiv:2310.02124.
  64. 64.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022a. Opt: Open pre-trained transformer language models. arXiv e-prints, pages arXiv–2205.
  65. 65.Yifan Zhang, Jingqin Yang, Yang Yuan, and Andrew Chi-Chih Yao. 2023b. Cumulative reasoning with large language models. arXiv preprint arXiv:2308.04371.
  66. 66.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022b. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.

Citation

MLA
Wang, Q., et al. “Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6106–31, https://doi.org/10.18653/v1/2024.acl-long.331.
APA
Wang, Q., Wang, Z., Su, Y., Tong, H., & Song, Y. (2024). Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6106–6131. https://doi.org/10.18653/v1/2024.acl-long.331
Chicago
Wang, Q., Z. Wang, Y. Su, H. Tong, and Y. Song. 2024. “Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6106–31. https://doi.org/10.18653/v1/2024.acl-long.331.
Harvard
Wang, Q. et al. (2024) “Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6106–6131. Available at: https://doi.org/10.18653/v1/2024.acl-long.331.
Vancouver
1. Wang Q, Wang Z, Su Y, Tong H, Song Y (2024) Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6106–6131

BibTeX

@inproceedings{wang-etal-2024-rethinking-bounds,
    title = "Rethinking the Bounds of {LLM} Reasoning: Are Multi-Agent Discussions the Key?",
    author = "Wang, Qineng  and
      Wang, Zihao  and
      Su, Ying  and
      Tong, Hanghang  and
      Song, Yangqiu",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.331/",
    doi = "10.18653/v1/2024.acl-long.331",
    pages = "6106--6131"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/