Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Tian LiangZhiwei HeWenxiang JiaoXing WangYan WangRui WangYujiu YangShuming ShiZhaopeng Tu

article2024EMNLP1,415 citations

Proposes a multi-agent debate framework that overcomes the cognitive stagnation of single-model self-reflection by using competing agents and an impartial judge to stimulate divergent thinking in complex reasoning tasks.

Listen

Large language models frequently struggle with complex reasoning and nuanced language tasks where initial, intuitive answers are misleading. Common correction techniques, such as self-reflection—where a model reviews and refines its own output—often fail due to the Degeneration-of-Thought problem. In this failure mode, once a model establishes confidence in an initial incorrect stance, it becomes rigid and fails to generate novel ideas or correct itself. This creates significant risks when deploying language models in high-stakes environments that require rigorous problem-solving.

The main objective of the article is to introduce and evaluate the Multi-Agent Debate framework, a collaborative problem-solving approach designed to overcome Degeneration-of-Thought by simulating multi-agent debates managed by an independent judge.

To evaluate this approach, the researchers tested the debate architecture against standard zero-shot baselines, chain-of-thought prompting, and self-reflection techniques across multiple language models, including GPT-3.5-Turbo, GPT-4, and open-source Vicuna variants. The evaluation focused on two demanding benchmarks: Commonsense Machine Translation (1,000 Chinese-to-English examples featuring lexical and syntactic ambiguities) and Counter-Intuitive Arithmetic Reasoning (200 problems containing intuitive traps requiring multi-step logic). Performance was evaluated using standard automated metrics, direct human assessments, and text diversity scores.

The findings show that the Multi-Agent Debate framework consistently outperforms existing reflection methods. On the counter-intuitive math dataset, the framework improved accuracy from 26.0% (standard GPT-3.5) and 27.5% (self-reflection) to 37.0%, representing an approximate 35% relative gain over self-reflection. On commonsense translation, GPT-3.5 augmented with the debate framework outperformed baseline GPT-4 across both automated scores and professional human ratings. Furthermore, the debate setup increased candidate text diversity from 19.3 to 49.7 and decreased translation bias from 29.0 to 24.8 compared to self-reflection. The analysis revealed that debate performance depends heavily on strong debater models rather than the judge model, that a moderate level of opposition yields better results than forced total disagreement, and that limiting debates to two agents avoids performance degradation caused by long-context confusion.

These results indicate that externalized peer feedback between agents effectively breaks the cognitive rigidity inherent in single-model reflection. For enterprise applications, this framework provides a viable pathway to achieve high-tier performance from smaller or less expensive base models on complex analytical tasks. However, this accuracy gain incurs a higher computational cost: the debate framework generates roughly 2.46 times the tokens of standard baseline prompting, representing a moderate increase over self-reflection's 1.83 times multiplier.

Organizations deploying language models for complex decision support should consider multi-agent debate structures with an adaptive stopping mechanism rather than relying on standard self-reflection. Implementers should maintain a balanced, moderate level of debate tension and ensure that the judge and debaters either share the same model architecture or use explicitly separated designs to avoid model-preference biases. Further engineering is recommended to enhance long-text processing before scaling debate configurations beyond two debaters.

The article's conclusions are supported by structured empirical testing across multiple benchmarks and human ratings. Readers should note that the framework was evaluated primarily on translation and arithmetic tasks with up to three debate iterations; confidence remains high for structured reasoning tasks, but further testing is needed to confirm generalizability across broader enterprise domains.

arXiv: 2305.19118
Cover for Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Abstract

Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies. Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively. However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect. To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution. Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation. Experiment results on two challenging datasets, commonsense machine translation and counter-intuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework. Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance. Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Multi-Agent Debate Framework
  • 3 Experiment
  • 3.1 Challenging Testbeds
  • 3.2 Setups
  • 3.3 Results on Common MT
  • 3.4 Results on Counter-Intuitive AR
  • 4 Analysis
  • 4.1 Mitigation of DoT
  • 4.2 Analysis of Judge
  • 4.3 Analysis of Debaters
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Challenging Testbeds
  • A.1 Commonsense Machine Translation
  • A.2 Counter-Intuitive Arithmetic Reasoning
  • B Human Evaluation Details
  • C Results on math and symbolic reasoning tasks
  • D Prompts for Different Debate Levels
  • E Extra Computational Cost
  • F Debate Process
  • F.1 Commonsense Machine Translation
  • F.2 Counter-Intuitive Arithmetic Reasoning

Knowls

  1. Knowl 1 — Degeneration-of-Thought in LLM Self-Reflection

    definition

    The Degeneration-of-Thought (DoT) problem describes a failure mode in large language model (LLM) self-reflection where, once an LLM-based agent establishes confidence in its generated answer or chain-of-thought, it fails to generate novel thoughts or alternative perspectives in subsequent self-reflection iterations, even if its initial answer or stance is incorrect.

    DoT is driven by three primary cognitive factors in LLMs:

    1. Bias and Distorted Perception: Pre-trained distributions contain inherent biases and heuristics that lead models instinctively to incorrect conclusions that persist during internal introspection.
    2. Rigidity and Resistance to Change: When self-reflecting without external prompts to counter its assumptions, the LLM tends to validate its own prior reasoning rather than challenge it.
    3. Limited External Feedback: Standard self-reflection is purely internal; lacking external counterarguments or alternative viewpoints, the model cannot identify its own blind spots.
  2. Knowl 2 — Multi-Agent Debate Framework

    model/method

    The Multi-Agent Debate (MAD) framework organizes multiple LLM agents to resolve complex reasoning or generation tasks through structured debate, counteracting the Degeneration-of-Thought problem via divergent thinking.

    The framework consists of:

    1. Meta Prompts: System instructions defining the topic, the number of participants, iteration bounds, and a "tit for tat" behavioral directive requiring agents to question and debate opposing views rather than achieve premature consensus.
    2. Debaters: A set of NN debater agents D={Di}i=1ND = \{D_i\}_{i=1}^N. In each iteration, debater DiD_i speaks sequentially based on the accumulated debate history HH, producing argument Di(H)=hD_i(H) = h. Typically, debaters are divided into affirmative and negative roles, where the negative agent actively challenges the affirmative agent's claims.
    3. Judge Agent: A separate LLM agent JJ that monitors and controls the debate using two operating modes:
      • Discriminative Mode (JdJ_d): At the conclusion of each debate round, the judge evaluates the history HH to decide whether a verified correct solution has emerged: Jd(H)={True,solution obtainedFalse,otherwiseJ_d(H) = \begin{cases} \text{True}, & \text{solution obtained} \\ \text{False}, & \text{otherwise} \end{cases} If Jd(H)=TrueJ_d(H) = \text{True}, the debate terminates immediately (adaptive break). If False\text{False}, the debate proceeds to the next round.
      • Extractive Mode (JeJ_e): If the debate reaches the maximum iteration limit without early termination, the judge extracts the final answer from the entire debate record: a=Je(H)a = J_e(H).
  3. Knowl 3 — Performance of Multi-Agent Debate on Commonsense Machine Translation

    empirical result

    On the Commonsense Machine Translation (Common MT) benchmark, Chinese-to-English translation performance was evaluated across three ambiguity categories: Lexical ambiguity, Contextless syntactic ambiguity, and Contextual syntactic ambiguity. Evaluation was conducted using COMET, BLEURT, and human assessment (HUMAN, scored from 1 to 5, Krippendorff's α=0.76\alpha = 0.76).

    Method Lexical Contextless Contextual
    COMET BLEURT HUMAN COMET BLEURT HUMAN COMET BLEURT HUMAN
    GPT-4 82.0 70.1 3.41 84.7 73.6 3.63 85.0 73.7 3.65
    GPT-3.5-Turbo 80.3 68.2 3.14 84.0 72.9 3.43 84.9 73.4 3.57
    + Rerank 80.9 68.6 3.16 84.5 73.2 3.46 85.3 73.9 3.58
    + MAPS 81.9 70.1 3.43 84.2 73.5 3.45 85.2 74.0 3.56
    + Self-Reflect 81.0 69.1 3.43 83.6 72.2 3.46 84.9 73.5 3.63
    + MAD 82.0 70.9 3.78 84.8 73.7 3.67 85.3 74.0 3.67
    Vicuna-7b 74.9 62.0 2.55 78.3 64.6 2.53 80.2 68.2 3.23
    + MAD 75.6 62.6 2.67 78.6 66.0 2.69 81.8 69.9 3.27
    Vicuna-13b 76.6 63.7 2.81 77.6 66.8 3.04 82.2 70.0 3.37
    + MAD 77.2 65.1 2.96 80.1 67.3 3.11 82.6 70.9 3.45

    Applying the Multi-Agent Debate (MAD) framework to GPT-3.5-Turbo achieves higher performance than standalone GPT-4 on Lexical ambiguity (COMET 82.0 vs. 82.0, BLEURT 70.9 vs. 70.1, HUMAN 3.78 vs. 3.41), Contextless ambiguity (COMET 84.8 vs. 84.7, BLEURT 73.7 vs. 73.6, HUMAN 3.67 vs. 3.63), and Contextual ambiguity (COMET 85.3 vs. 85.0, BLEURT 74.0 vs. 73.7, HUMAN 3.67 vs. 3.65). MAD consistently outperforms standard iterative Self-Reflection across all backbone models.

  4. Knowl 4 — Performance of Multi-Agent Debate on Counter-Intuitive and Multi-Step Reasoning

    empirical result

    On reasoning tasks where initial intuitive impressions lead to incorrect answers, Multi-Agent Debate (MAD) demonstrates substantial accuracy gains over standard prompting, chain-of-thought (CoT), self-consistency, and self-reflection.

    On the Counter-Intuitive Arithmetic Reasoning (CIAR) benchmark (200 questions evaluated by accuracy ACC):

    Method ACC (%)
    GPT-4 51.0
    GPT-3.5-Turbo 26.0
    + CoT 28.0
    + Self-Consistency 29.5
    + Self-Reflect 27.5
    + MAD 37.0

    On standard mathematical and symbolic reasoning datasets (using GPT-3.5-Turbo):

    Method Math Reasoning Symbolic Reasoning (BBH)
    GSM AddSub Penguin Date Colored Objects
    CoT 70.2 87.3 58.9 56.4 57.2
    Self-Reflect 70.8 87.6 61.0 58.0 58.0
    MAD 73.8 92.1 63.7 65.2 58.8

    On CIAR, Self-Reflect improves baseline GPT-3.5-Turbo accuracy by only 1.5% (from 26.0% to 27.5%), whereas MAD improves accuracy by 11.0% (reaching 37.0%).

  5. Knowl 5 — Counter-Intuitive Arithmetic Reasoning Dataset

    experimental setup

    The Counter-Intuitive Arithmetic Reasoning (CIAR) benchmark consists of 200 mathematical reasoning questions specifically constructed to evaluate an LLM's capacity for deep, deliberate contemplation versus reliance on fast superficial heuristics.

    The benchmark exhibits two main characteristics:

    1. Resistance to Intuition: Questions feature deceptive surface structures designed to trigger appealing but incorrect answers when solved by intuitive heuristics (e.g., averaging speeds over equal distances rather than calculating harmonic means).
    2. Multi-Step Reasoning: Correctly answering each problem requires a non-trivial multi-step deduction process.

    Each dataset entry includes:

    • The target question,
    • The ground-truth correct answer and a step-by-step mathematical explanation,
    • A common intuitive but incorrect answer and the false reasoning chain leading to it.
  6. Knowl 6 — Adaptive Break Strategy in Multi-Agent Debate

    empirical result

    In the Multi-Agent Debate (MAD) framework, the judge agent dynamically evaluates debate history after each round to execute an early adaptive break when an optimal answer is obtained, rather than running for a fixed number of rounds.

    Empirical evaluations demonstrate:

    1. Iteration Distribution: In the majority of instances (~130 out of 200 samples on Common MT), the correct translation is reached and verified in iteration 1. Harder examples (corresponding to lower human quality subsets) require 2 or 3 iterations before the judge obtains adequate consensus.
    2. Fixed-Round Degradation: When MAD is forced to run for fixed iteration counts without adaptive early termination, translation quality on Common MT peaks at iteration 1 (COMET score ~81.3) and monotonically degrades with subsequent rounds (dropping to ~80.3 at iteration 5). In contrast, the dynamic adaptive break strategy achieves an overall COMET score of 82.0 by terminating unproblematic examples early while allowing complex examples to continue.
  7. Knowl 7 — Impact of Disagreement Intensity on Multi-Agent Debate

    empirical result

    The level of opposition ("tit for tat" behavior) enforced via debater meta prompts directly governs debate effectiveness and task performance.

    Four discrete debate instructions were tested on Common MT Lexical ambiguity resolution:

    • Level 0 (Full Consensus Required): "Both sides must reach a full consensus on every point of the debate..." -> Average disagreement: 0.00, Ambiguity Resolution: 0.63.
    • Level 1 (Minor Consensus Allowed): "Most of the debate should be characterized by disagreements, but there may still be a small amount of consensus on less significant points." -> Average disagreement: 0.35, Ambiguity Resolution: 0.66.
    • Level 2 (Default - Modest Opposition): "It’s not necessary to fully agree with each other’s perspectives, as our objective is to find the correct answer." -> Average disagreement: 0.58, Ambiguity Resolution: 0.80.
    • Level 3 (Forced Total Disagreement): "Both sides must disagree with each other on every point of the debate. There should be no consensus whatsoever." -> Average disagreement: 0.99, Ambiguity Resolution: 0.72.

    A moderate level of opposition (Level 2) achieves the best performance. Enforcing total disagreement (Level 3) induces argument polarization, where agents prioritize winning arguments over finding the truth, which harms final accuracy.

  8. Knowl 8 — Debater Dominance and Ingroup Preference Bias in Judge LLMs

    empirical result

    Analysis of agent assignments in the Multi-Agent Debate framework reveals two key phenomena regarding judge and debater roles:

    1. Debater Quality Dominates Judge Quality: Setting GPT-3.5-Turbo as debaters with a smaller model (Vicuna-13b) as judge achieves COMET 83.2 / HUMAN 3.47 on Common MT, substantially higher than setting Vicuna-13b as debaters with GPT-3.5-Turbo as judge (COMET 80.4 / HUMAN 3.25). The debater models define the performance ceiling of the debate.
    2. Judge Ingroup LLM Preference: When debaters are instantiated with different LLM architectures, an LLM judge exhibits an ingroup preference toward the debater sharing its own backbone model:
    Judge Affirmative Negative Affirmative Wins Negative Wins
    GPT-4 GPT-3.5-Turbo GPT-4 52 136
    GPT-4 GPT-4 GPT-3.5-Turbo 120 77
    GPT-4 GPT-4 GPT-4 67 124
    GPT-3.5-Turbo GPT-3.5-Turbo GPT-3.5-Turbo 87 104

    When all agents share the same backbone LLM, judges consistently show a functional preference for the negative side (e.g., 124 negative vs. 67 affirmative wins under GPT-4), which provides corrective refinement to initial claims.

  9. Knowl 9 — Debater Scaling Bottleneck and Inference Cost in Multi-Agent Debate

    limitation

    The Multi-Agent Debate (MAD) framework has two operational limitations:

    1. Performance Drop with Additional Debaters: Increasing the debater count beyond two reduces performance on Common MT:

      • 2 Debaters: COMET 84.4, HUMAN 3.69
      • 3 Debaters: COMET 83.1, HUMAN 3.58
      • 4 Debaters: COMET 82.9, HUMAN 3.49

      This degradation occurs because increasing debater count lengthens the context window, causing LLM debaters to forget prior arguments and increasing the difficulty for the judge to extract and summarize key points.

    2. Computational Overhead: Relative to zero-shot Chain-of-Thought (1.0×1.0\times generated token cost), Self-Reflection incurs 1.83×1.83\times generated tokens, and MAD incurs 2.46×2.46\times generated tokens.

  10. Knowl 10 — Ambiguity Bias Reduction and Diversity Enhancement in Thought Generation

    empirical result

    The mitigation of Degeneration-of-Thought (DoT) by Multi-Agent Debate was quantified on Common MT by measuring translation ambiguity error rate (Bias) and generation diversity (Diversity).

    Diversity between the initial translation candidate Cand1\text{Cand}_1 (the affirmative response) and the subsequent translation candidate Cand2\text{Cand}_2 (the negative response) is computed as:

    Diversity=100−Self_BLEU(Cand1,Cand2)\text{Diversity} = 100 - \text{Self\_BLEU}(\text{Cand}_1, \text{Cand}_2)
    Method Bias (%) ↓\downarrow Diversity ↑\uparrow
    Self-Reflect 29.0 19.3
    MAD 24.8 49.7

    MAD reduces ambiguity bias from 29.0% to 24.8% and increases candidate diversity from 19.3 to 49.7 compared to Self-Reflection, demonstrating that multi-agent confrontation forces the exploration of non-redundant reasoning paths.

Coverage note — None was omitted; all contributed definitions, framework components, benchmark specifications, primary empirical results, ablation analyses (debate levels, judge biases, debater scaling), and limitations are included.

References

  1. 1.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 675–718.
  2. 2.Lisa Bortolotti. 2011. Does reflection lead to wise choices? Philosophical Explorations, 14(3):297–313.
  3. 3.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  4. 4.Kahneman Daniel. 2017. Thinking, fast and slow. Farrar, Straus and Giroux.
  5. 5.Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang. 2023. Active prompting with chain-of-thought for large language models. arXiv preprint arXiv:2302.12246.
  6. 6.Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023. Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325.
  7. 7.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. 2023. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142.
  8. 8.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  9. 9.Xavier Garcia, Yamini Bansal, Colin Cherry, George Foster, Maxim Krikun, Melvin Johnson, and Orhan Firat. 2023. The unreasonable effectiveness of few-shot learning for machine translation. In International Conference on Machine Learning, pages 10867–10878. PMLR.
  10. 10.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023. Critic: Large language models can self-correct with tool-interactive critiquing.
  11. 11.Jie He, Tao Wang, Deyi Xiong, and Qun Liu. 2020. The box is in the pen: Evaluating commonsense reasoning in neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3662–3672, Online. Association for Computational Linguistics.
  12. 12.Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang. 2024. Exploring human-like translation strategy with large language models. Transactions of the Association for Computational Linguistics, 12:229–246.
  13. 13.Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210.
  14. 14.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 523–533.
  15. 15.Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. Is chatgpt a good translator? yes with gpt-4 as the engine. arXiv preprint arXiv:2301.08745.
  16. 16.Machiel Keestra. 2017. Metacognition and reflection by interdisciplinary experts: Insights from cognitive science and philosophy. Issues in Interdisciplinary Studies, 35:121–169.
  17. 17.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  18. 18.Yuqing Kong, Yunqi Li, Yubo Zhang, Zhihuan Huang, and Jinzhao Wu. 2022. Eliciting thinking hierarchy without a prior. Advances in Neural Information Processing Systems, 35:13329–13341.
  19. 19.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
  20. 20.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36.
  21. 21.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pages 1–22.
  22. 22.Jonathan Pilault, Xavier Garcia, Arthur Bražinskas, and Orhan Firat. 2023. Interactive-chain-prompting: Ambiguity resolution for crosslingual conditional generation with interaction. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 455–483.
  23. 23.Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1743–1752.
  24. 24.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
  25. 25.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research.
  26. 26.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051.
  27. 27.Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926.
  28. 28.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  29. 29.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  30. 30.Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael Lyu. 2023. Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark. arXiv preprint arXiv:2303.13648.
  31. 31.Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, and Bing Qin. 2023. Diving into the inter-consistency of large language models: An insightful analysis through debate. arXiv preprint arXiv:2305.11595.
  32. 32.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36.
  33. 33.Haiyan Yin, Dingcheng Li, Xu Li, and Ping Li. 2020. Meta-cotgan: A meta cooperative training paradigm for improving adversarial text generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9466–9473.
  34. 34.Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant. 2023. Answering questions by meta-reasoning over multiple chains of thought. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5942–5966.
  35. 35.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  36. 36.Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li. 2023. Progressive-hint prompting improves reasoning in large language models. arXiv preprint arXiv:2304.09797.
  37. 37.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Jiaxing Zhang, Yujiu Yang, et al. 2023a. Solving math word problems via cooperative reasoning induced language models. In The 61st Annual Meeting Of The Association For Computational Linguistics.
  38. 38.Xinyu Zhu, Cheng Yang, Bei Chen, Siheng Li, Jian-Guang Lou, and Yujiu Yang. 2023b. Question answering as programming for solving time-sensitive questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12775–12790.
  39. 39.Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al. 2023c. Ghost in the minecraft: Generally capable agents for open-world enviroments via large language models with text-based knowledge and memory. arXiv preprint arXiv:2305.17144.

Citation

MLA
Liang, T., et al. “Encouraging Divergent Thinking in Large Language Models Through Multi-Agent Debate”. arXiv, 2023, http://arxiv.org/abs/2305.19118v4.
APA
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., & Tu, Z. (2023). Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. arXiv. http://arxiv.org/abs/2305.19118v4
Chicago
Liang, T., Z. He, W. Jiao, et al. 2023. “Encouraging Divergent Thinking in Large Language Models Through Multi-Agent Debate”. arXiv. http://arxiv.org/abs/2305.19118v4.
Harvard
Liang, T. et al. (2023) “Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.19118v4.
Vancouver
1. Liang T, He Z, Jiao W, Wang X, Wang Y, Wang R, Yang Y, Shi S, Tu Z (2023) Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. arXiv

BibTeX

@article{liang2023encouraging,
  title = {Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate},
  author = {Liang, Tian and He, Zhiwei and Jiao, Wenxiang and Wang, Xing and Wang, Yan and Wang, Rui and Yang, Yujiu and Shi, Shuming and Tu, Zhaopeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.19118v4},
  eprint = {2305.19118}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/