Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication

Zhangyue YinQiushi SunCheng ChangQipeng GuoJunqi DaiXuanjing HuangXipeng Qiu

article2023EMNLP113 citations

Proposes Exchange-of-Thought, a collaborative framework using network-inspired communication paradigms and confidence evaluation to let multiple large language models share reasoning steps and improve complex problem-solving accuracy cost-effectively.

Listen

Large language models often struggle with complex, multi-step reasoning tasks. While existing techniques such as chain-of-thought prompting and iterative self-correction help guide models through intermediate steps, they rely entirely on a single model's internal knowledge and perspectives. When a model makes an initial misstep, it frequently lacks the external feedback needed to recognize and rectify its errors.

The article introduces and evaluates Exchange-of-Thought (EoT), a framework designed to improve reasoning accuracy by enabling multiple language models to communicate, share intermediate reasoning steps, and critique one another during problem-solving.

To evaluate this framework, the authors tested communication architectures inspired by network structures across mathematical, commonsense, and symbolic reasoning benchmarks. These setups included Memory (a fully visible shared logbook), Report (a centralized hub), Relay (a circular chain), and Debate (a hierarchical tree). The evaluation incorporated a confidence-tracking mechanism—which measures a model's certainty based on the stability of its answers across rounds—and a consensus-based stopping rule. The primary experiments deployed three GPT-3.5 models communicating collaboratively and benchmarked them against single-model prompting, majority-voting ensembles, progressive hint methods, and standalone GPT-4.

The article demonstrates that cross-model communication delivers significant performance and efficiency advantages. First, across all four communication structures, the framework improved reasoning accuracy over standard chain-of-thought and established baselines, boosting mathematical reasoning accuracy by roughly 3.0 to 3.3 percentage points over progressive prompting. Second, collaborative interaction among three smaller GPT-3.5 models achieved results comparable to, and in some datasets higher than, a single, much larger GPT-4 model. Third, the framework delivered these improvements with lower computational costs—reducing token expenses by 20% compared to a five-sample majority-voting baseline while increasing accuracy by about 3%, and matching the performance of complex ten-path sampling at approximately one-seventh of the cost. Fourth, confidence scoring proved essential, lifting accuracy by an average of 2.92 percentage points by preventing the propagation of erroneous reasoning chains. Finally, performance further improved when mixing distinct model families, such as combining GPT-3.5, GPT-4, and Claude-2.

These findings indicate that collaborative multi-model architectures can reduce dependency on massive, expensive proprietary systems. Organizations can lower operational costs, improve output reliability, and decrease computational overhead by orchestrating smaller or heterogeneous models into structured communication workflows rather than relying solely on larger models or repetitive single-model sampling.

Organizations developing complex reasoning systems should consider adopting multi-model communication paradigms, selecting topologies that align with their operational needs. For example, centralized Report architectures suit setups with one high-capacity model, whereas decentralized Relay or Memory paradigms fit homogeneous systems. Teams should also implement stability-based confidence metrics to filter out early errors and pilot heterogeneous model mixtures to maximize diverse reasoning.

These results are subject to certain boundaries. The experiments evaluated proprietary commercial models and restricted group sizes to three participants to manage token context limits and operational costs. Open-source models and larger model groups were not evaluated. While confidence in the tested commercial benchmarks is high, decision-makers should conduct pilot validations before deploying multi-agent communication frameworks in large-scale production environments.

Cover for Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication

Abstract

Large Language Models (LLMs) have recently made significant strides in complex reasoning tasks through the Chain-of-Thought technique. Despite this progress, their reasoning is often constrained by their intrinsic understanding, lacking external insights. To address this, we propose Exchange-of-Thought (EoT), a novel framework that enables cross-model communication during problem-solving. Drawing inspiration from network topology, EoT integrates four unique communication paradigms: Memory, Report, Relay, and Debate. This paper delves into the communication dynamics and volume associated with each paradigm. To counterbalance the risks of incorrect reasoning chains, we implement a robust confidence evaluation mechanism within these communications. Our experiments across diverse complex reasoning tasks demonstrate that EoT significantly surpasses established baselines, underscoring the value of external insights in enhancing LLM performance. Furthermore, we show that EoT achieves these superior results in a cost-effective manner, marking a promising advancement for efficient and collaborative AI problem-solving.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Chain-of-Thought prompting in LLMs
  • 2.2 Ensemble of Reasoning Paths
  • 2.3 Reasoning Path Refinement
  • 3 Preliminary
  • 4 Methodology
  • 4.1 Communication Paradigm
  • 4.2 Communication Volume
  • 4.3 Termination Condition
  • 4.4 Confidence Evaluation
  • 5 Experiments
  • 5.1 Experimental Setups
  • 5.2 Performance of EoT
  • 5.3 Discussions
  • 6 Conclusion
  • Ethics Statement
  • Acknowledgement
  • References
  • A Limitations and Broader Impacts
  • B Datasets and Evaluation Metrics
  • C Implementation Details
  • E Case Studies

Knowls

  1. Knowl 1 — Exchange-of-Thought framework

    model/method

    Exchange-of-Thought (EoT) is a framework in which multiple large language models solve the same problem while exchanging complete reasoning chains and answers, rather than communicating only final answers. Let qq be a question, DD be any demonstrations included in the prompt, and M={m1,…,mn}M=\{m_1,\ldots,m_n\} be nn participating models, where model mim_i has parameters θi\theta_i. In communication round jj, model mim_i produces a rationale ri(j)r_i^{(j)} and answer ai(j)a_i^{(j)} using the previous-round rationale-answer pairs made available to it by the communication topology, indexed by KiK_i:

    (ri(j),ai(j))∼Pθi(ri,ai∣D,q,{(rℓ(j−1),aℓ(j−1)):ℓ∈Ki}).(r_i^{(j)},a_i^{(j)})\sim P_{\theta_i}\left(r_i,a_i\mid D,q,\{(r_\ell^{(j-1)},a_\ell^{(j-1)}):\ell\in K_i\}\right).

    In the first round, no exchanged reasoning is available, so every model independently generates a chain-of-thought solution conditioned on DD and qq. Subsequent rounds prompt each model to reconsider its previous solution in light of the reasoning received from other models. Communication continues until a termination rule is satisfied, and the final answer is selected from the participating models’ answers, using majority selection in the paper’s default implementation. EoT therefore supplies external model perspectives that ordinary chain-of-thought and self-correction do not provide.

  2. Knowl 2 — Four cross-model communication paradigms

    model/method

    EoT defines four ways to connect models, each corresponding to a distinct information-flow topology:

    • Memory: Every model writes its rationale and answer to a shared logbook visible to all models. Each model can inspect every other model’s previous-round output. This maximizes information availability and speed of dissemination, but has the highest communication volume.

    • Report: One model is designated as a central node. It receives the previous-round outputs of all other models, while every noncentral model receives information only from the central model. This gives rapid information flow but concentrates processing and analysis in the central model.

    • Relay: Models are arranged in a directed ring. Each model sends its current rationale and answer to the next model and receives the previous-round output from the preceding model. This distributes the processing burden across models, but information can take several rounds to propagate around the ring.

    • Debate: Models are arranged as a binary tree. Child models can exchange information, while parent models aggregate information flowing upward from their children. This balances the processing burden and propagation speed: parent nodes synthesize child reasoning, while leaf nodes provide alternative solutions.

  3. Knowl 3 — Communication-volume and propagation analysis

    theoretical result

    For nn communicating models, the paper measures communication volume as the number of messages received in one round, with each model transmitting its previous-round rationale and answer according to its topology. The resulting volumes and propagation properties are:

    • Memory has volume V=n2V=n^2; each piece of information reaches a destination in one transmission.
    • Report has volume V=3n−2V=3n-2: the central model processes nn messages and each of the n−1n-1 noncentral models processes two. Information requires at most two transmissions to pass through the central node.
    • Relay has volume V=2nV=2n. A model receives its own previous-round information and the information from its predecessor. Information from the next model requires n−1n-1 transmissions to reach it around the ring, giving an average propagation distance of (n−1)/2(n-1)/2 transmissions.
    • Debate uses a binary tree of height h=⌈log⁡2(n+1)⌉h=\lceil\log_2(n+1)\rceil. Under the paper’s binary-tree accounting and the stated assumption that nn is odd, its communication volume is V=72(n−1)V=\frac{7}{2}(n-1). Aggregating information between two nodes requires at most 2(h−1)2(h-1) transmissions when other communications are ignored.

    Thus, Memory provides the most globally visible and rapidly disseminated information but costs the most messages, Report provides fast centralized aggregation, Relay minimizes centralized processing at the expense of propagation speed, and Debate provides an intermediate tree-structured trade-off.

  4. Knowl 4 — Termination and confidence control

    model/method

    EoT uses two termination rules and a confidence score to limit unnecessary exchanges and reduce the influence of erroneous reasoning.

    Termination. Under consistent-output termination, model mim_i stops sending and receiving information when its current answer equals its answer from the immediately preceding round, ai(j)=ai(j−1)a_i^{(j)}=a_i^{(j-1)}. This is an individual stopping rule. Under majority-consensus termination, communication stops globally once a majority of models produce the same answer. In the three-model implementation, unanimous agreement is required during the first five rounds to avoid prematurely accepting a two-model agreement; if unanimity has not occurred by round five, the majority answer is selected. If a model exits under the individual rule, the implementation terminates the three-model interaction and uses the exiting model’s answer.

    Confidence evaluation. Suppose model mim_i has generated answers ai(1),…,ai(k)a_i^{(1)},\ldots,a_i^{(k)} over kk rounds. Let cic_i be the number of occurrences of the most frequent answer among these kk answers. Its confidence is defined as

    Ci=cik,0≤Ci≤1.C_i=\frac{c_i}{k},\qquad 0\leq C_i\leq 1.

    Confidence information is introduced from the second communication round by prepending the confidence score to the model’s proposed solution. A stable answer receives a high score, while frequent answer changes receive a low score. The communication prompt warns that a low confidence score below 0.50.5 suggests a high probability of error, while also noting that high-confidence answers can still be wrong.

  5. Knowl 5 — Experimental tasks and evaluation protocol

    experimental setup

    EoT was evaluated on mathematical, commonsense, and symbolic reasoning. The experiments used GPT-3.5-Turbo-0301 and GPT-4-0314 through the OpenAI API, with generation temperature set to 11. The default EoT configuration used three GPT-3.5-Turbo-0301 models, majority-consensus termination, confidence evaluation, and majority-answer selection. Results averaged five runs and report standard deviations where applicable. Baselines were chain-of-thought (CoT), complexity-based prompting (ComplexCoT), self-consistency applied to sampled chains (CoT-SC), and progressive-hint prompting (PHP). Accuracy was computed by extracting numerical answers for number-answer datasets and matching selected options for multiple-choice or true/false datasets.

    The evaluated datasets and test sizes were:

    Could not parse LaTeX table

    For auxiliary analyses subject to rate and cost limits, at most 1,000 samples were used per run. The paper’s cost calculations use GPT-3.5-Turbo-0301 pricing: 0.00150.0015 per 1,000 input tokens plus 0.0020.002 per 1,000 output tokens.

  6. Knowl 6 — Mathematical-reasoning performance

    data/table

    On six mathematical reasoning datasets, EoT consistently outperformed the single-chain and ensemble baselines under the same GPT-3.5 experimental setting. The table reports accuracy percentages; entries with ±\pm show the standard deviation across five runs. The EoT rows use three communicating GPT-3.5 models, while the GPT-4 CoT row is a separately sourced single-model result. EoT’s average accuracies were 87.20–87.50, exceeding PHP at 84.16 and CoT-SC(10) at 86.59. Relative to PHP, the average gains were 3.04 percentage points for Memory, 3.30 for Report, 3.34 for Relay, and 3.20 for Debate.

    Could not parse LaTeX table

    Three GPT-3.5 models with EoT surpassed a single GPT-4 with CoT on MultiArith and SingleEQ, although GPT-4 remained stronger on the other listed mathematical datasets.

  7. Knowl 7 — Cross-domain reasoning gains

    empirical result

    EoT improved performance beyond chain-of-thought and sampled-chain ensembles on commonsense and symbolic reasoning tasks. On StrategyQA, relative to CoT, the accuracy gains were 8.06 percentage points for EoT-Memory, 8.24 for EoT-Report, 8.42 for EoT-Relay, and 8.67 for EoT-Debate. Comparable gains were observed on CommonsenseQA, and all four EoT paradigms outperformed CoT-SC(10) on both commonsense datasets.

    On the Penguins in a Table symbolic dataset, relative to CoT-SC(3), the gains were 2.01 percentage points for Memory, 1.92 for Report, 2.33 for Relay, and 2.05 for Debate. On Date Understanding, all four EoT variants improved average accuracy by 2.1 percentage points over CoT-SC(10). These results indicate that the benefit of exchanging external reasoning is not restricted to arithmetic word problems.

  8. Knowl 8 — Termination and confidence ablations

    empirical result

    The paper’s ablations show that collective termination and confidence-aware communication materially affect EoT accuracy. On AQuA, replacing consistent-output termination with majority-consensus termination improved accuracy by 4.33 percentage points for Memory, 4.01 for Report, 7.56 for Relay, and 4.97 for Debate. The authors attribute the poorer individual stopping results partly to premature exits caused by output degeneration, whereas majority consensus allows collective negotiation.

    On GSM8K, adding confidence evaluation improved accuracy by an average of 2.92 percentage points across the four communication paradigms. The paper reports that confidence information helps models decide when to accept another model’s reasoning and reduces interference from incorrect chains.

    On SVAMP, most examples reached an answer consensus within three communication rounds for every paradigm. Only a minority of difficult examples required more than five rounds, so EoT generally terminates after a short exchange while allowing additional discussion when consensus is difficult.

  9. Knowl 9 — Cost, model applicability, and node-position effects

    empirical result

    EoT improved the accuracy-cost trade-off on GSM8K. Compared with CoT-SC(5), EoT reduced computational cost by 20% while increasing performance by 3%. It achieved performance comparable to ComplexCoT-SC(10) at approximately one-seventh of that method’s cost. The low cost was associated with the observation that most examples terminated within three communication rounds.

    EoT also transferred across model backbones. Relative to CoT-SC(5), EoT improved performance by 3.2 percentage points with GPT-3.5, 1.0 with GPT-4, and 1.4 with Claude-2. Model placement mattered in centralized topologies: placing GPT-4 at the central node in Report improved performance by more than 1% relative to placing it at a noncentral node, and placing GPT-4 at a parent node in Debate improved performance by 0.9% relative to placing it at a child node. GPT-4 placement had negligible effect in decentralized Memory and Relay. Configurations containing two GPT-4 models and one GPT-3.5 model substantially outperformed configurations with two GPT-3.5 models and one GPT-4 model, while a heterogeneous GPT-3.5/GPT-4/Claude-2 configuration performed close to or better than two GPT-4 models plus one GPT-3.5 model, suggesting that model diversity can be beneficial.

  10. Knowl 10 — Scope limitations of the evaluation

    limitation

    The experiments did not include open-source language models because the authors considered their communication and analytical capabilities, as well as their computational requirements, insufficient for the study’s setting. The claim that capable open-source models could match or exceed commercial models through exchanged insights is presented as a possibility rather than a demonstrated result.

    The number of communicating models was also kept small because exchanging long reasoning texts is constrained by current context windows. The paper therefore does not establish how EoT scales to larger groups or substantially longer communication histories. The authors identify long-context language-model techniques as a prerequisite for broader multi-model communication.

Coverage note — The pilot error-analysis examples and detailed role-playing prompts/case studies were omitted because they illustrate the framework but add little load-bearing content beyond the formal method and empirical results.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv preprint, abs/2204.05862.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022b. Constitutional ai: Harmlessness from ai feedback.
  3. 3.BIG bench authors. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research.
  4. 4.Nivedita Bisht and Sapna Singh. 2015. Analytical study of different network topologies. International Research Journal of Engineering and Technology (IRJET), 2(01):88–90.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation.
  7. 7.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. ArXiv preprint, abs/2211.12588.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  9. 9.Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. 2023. A survey of chain of thought reasoning: Advances, frontiers and future.
  10. 10.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. ArXiv preprint, abs/2210.11416.
  11. 11.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems.
  12. 12.Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang. 2023. Active prompting with chain-of-thought for large language models. ArXiv preprint, abs/2302.12246.
  13. 13.Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023. Improving factuality and reasoning in language models through multi-agent debate. ArXiv preprint, abs/2305.14325.
  14. 14.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. 2023a. Improving language model negotiation with self-play and in-context learning from ai feedback.
  15. 15.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2023b. Complexity-based prompting for multi-step reasoning. In The Eleventh International Conference on Learning Representations.
  16. 16.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. Pal: Program-aided language models. ArXiv preprint, abs/2211.10435.
  17. 17.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions on Computational Linguistics, 9:346–361.
  18. 18.David Ha and Yujin Tang. 2022. Collective intelligence for deep learning: A survey of recent developments. Collective Intelligence, 1(1):26339137221114874.
  19. 19.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 523–533. Association for Computational Linguistics.
  20. 20.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023. Large language models cannot self-correct reasoning yet.
  21. 21.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. ArXiv preprint, abs/2001.08361.
  22. 22.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  23. 23.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  24. 24.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics, 3:585–597.
  25. 25.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157, San Diego, California. Association for Computational Linguistics.
  26. 26.Ludmila I Kuncheva and Christopher J Whitaker. 2003. Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Machine learning, 51:181–207.
  27. 27.Gustave Le Bon. 1897. The crowd: A study of the popular mind. TF Unwin.
  28. 28.Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267.
  29. 29.Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu. 2023a. Unified demonstration retriever for in-context learning. ArXiv preprint, abs/2305.04320.
  30. 30.Xiaonan Li and Xipeng Qiu. 2023a. Finding supporting examples for in-context learning. ArXiv preprint, abs/2302.13539.
  31. 31.Xiaonan Li and Xipeng Qiu. 2023b. Mot: Pre-thinking and recalling enable chatgpt to self-improve with memory-of-thoughts. ArXiv preprint, abs/2305.05181.
  32. 32.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2023b. Making language models better reasoners with step-aware verifier. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5315–5333, Toronto, Canada. Association for Computational Linguistics.
  33. 33.Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023. Encouraging divergent thinking in large language models through multi-agent debate. ArXiv preprint, abs/2305.19118.
  34. 34.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 158–167. Association for Computational Linguistics.
  35. 35.Xiaoran Liu, Hang Yan, Shuo Zhang, Chenxin An, Xipeng Qiu, and Dahua Lin. 2023. Scaling laws of rope-based extrapolation.
  36. 36.Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023. Faithful chain-of-thought reasoning. ArXiv preprint, abs/2301.13379.
  37. 37.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. ArXiv preprint, abs/2303.17651.
  38. 38.OpenAI. 2023. Gpt-4 technical report.
  39. 39.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  40. 40.S. Parsons and Peter McBurney. 2003. Argumentation-based communication between agents. In Communication in Multiagent Systems.
  41. 41.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, Online. Association for Computational Linguistics.
  42. 42.Pragaash Ponnusamy, Alireza Ghias, Yi Yi, Benjamin Yao, Chenlei Guo, and Ruhi Sarikaya. 2022. Feedback-based self-learning in large-scale conversational ai agents. AI magazine, 42(4):43–56.
  43. 43.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2022. Scaling language models: Methods, analysis & insights from training gopher.
  44. 44.Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. Parallel context windows for large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6383–6402, Toronto, Canada. Association for Computational Linguistics.
  45. 45.Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pages 1743–1752. The Association for Computational Linguistics.
  46. 46.Noah Shinn, Beck Labash, and Ashwin Gopinath. 2023. Reflexion: an autonomous agent with dynamic memory and self-reflection. ArXiv preprint, abs/2303.11366.
  47. 47.Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati. 2023. Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems.
  48. 48.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A contrastive framework for neural text generation. In Advances in Neural Information Processing Systems.
  49. 49.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. 2023. Challenging BIG-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051, Toronto, Canada. Association for Computational Linguistics.
  50. 50.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  51. 51.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. ArXiv preprint, abs/2201.08239.
  52. 52.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971.
  53. 53.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  54. 54.Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłos. 2023. ´ Focused transformer: Contrastive training for context scaling.
  55. 55.Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023. Can large language models really improve by self-critiquing their own plans?
  56. 56.Aimee Van Wynsberghe. 2021. Sustainable ai: Ai for sustainability and the sustainability of ai. AI and Ethics, 1(3):213–218.
  57. 57.Jianing Wang, Qiushi Sun, Nuo Chen, Xiang Li, and Ming Gao. 2023a. Boosting language models reasoning with chain-of-knowledge prompting. ArXiv preprint, abs/2306.06427.
  58. 58.Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. 2023b. Augmenting language models with long-term memory.
  59. 59.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023c. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  60. 60.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research.
  61. 61.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  62. 62.Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2023. Generating sequences by learning to self-correct. In The Eleventh International Conference on Learning Representations.
  63. 63.Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. 2022. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813.
  64. 64.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023. Efficient streaming language models with attention sinks.
  65. 65.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
  66. 66.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don’t know? In Findings of the Association for Computational Linguistics: ACL 2023, pages 8653–8665, Toronto, Canada. Association for Computational Linguistics.
  67. 67.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. ArXiv preprint, abs/2205.01068.
  68. 68.Tianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel, and Joseph E. Gonzalez. 2023a. The wisdom of hindsight makes language models better instruction followers. In Proceedings of the 40th International Conference on Machine Learning, ICML’23.
  69. 69.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023b. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations.
  70. 70.Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li. 2023. Progressive-hint prompting improves reasoning in large language models. ArXiv preprint, abs/2304.09797.

Citation

MLA
Yin, Z., et al. “Exchange-of-Thought: Enhancing Large Language Model Capabilities Through Cross-Model Communication”. arXiv, 2023, http://arxiv.org/abs/2312.01823v1.
APA
Yin, Z., Sun, Q., Chang, C., Guo, Q., Dai, J., Huang, X., & Qiu, X. (2023). Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication. arXiv. http://arxiv.org/abs/2312.01823v1
Chicago
Yin, Z., Q. Sun, C. Chang, et al. 2023. “Exchange-of-Thought: Enhancing Large Language Model Capabilities Through Cross-Model Communication”. arXiv. http://arxiv.org/abs/2312.01823v1.
Harvard
Yin, Z. et al. (2023) “Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.01823v1.
Vancouver
1. Yin Z, Sun Q, Chang C, Guo Q, Dai J, Huang X, Qiu X (2023) Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication. arXiv

BibTeX

@article{yin2023exchange,
  title = {Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication},
  author = {Yin, Zhangyue and Sun, Qiushi and Chang, Cheng and Guo, Qipeng and Dai, Junqi and Huang, Xuanjing and Qiu, Xipeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.01823v1},
  eprint = {2312.01823}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/