Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication
Zhangyue YinQiushi SunCheng ChangQipeng GuoJunqi DaiXuanjing HuangXipeng Qiu
Proposes Exchange-of-Thought, a collaborative framework using network-inspired communication paradigms and confidence evaluation to let multiple large language models share reasoning steps and improve complex problem-solving accuracy cost-effectively.
Large language models often struggle with complex, multi-step reasoning tasks. While existing techniques such as chain-of-thought prompting and iterative self-correction help guide models through intermediate steps, they rely entirely on a single model's internal knowledge and perspectives. When a model makes an initial misstep, it frequently lacks the external feedback needed to recognize and rectify its errors.
The article introduces and evaluates Exchange-of-Thought (EoT), a framework designed to improve reasoning accuracy by enabling multiple language models to communicate, share intermediate reasoning steps, and critique one another during problem-solving.
To evaluate this framework, the authors tested communication architectures inspired by network structures across mathematical, commonsense, and symbolic reasoning benchmarks. These setups included Memory (a fully visible shared logbook), Report (a centralized hub), Relay (a circular chain), and Debate (a hierarchical tree). The evaluation incorporated a confidence-tracking mechanism—which measures a model's certainty based on the stability of its answers across rounds—and a consensus-based stopping rule. The primary experiments deployed three GPT-3.5 models communicating collaboratively and benchmarked them against single-model prompting, majority-voting ensembles, progressive hint methods, and standalone GPT-4.
The article demonstrates that cross-model communication delivers significant performance and efficiency advantages. First, across all four communication structures, the framework improved reasoning accuracy over standard chain-of-thought and established baselines, boosting mathematical reasoning accuracy by roughly 3.0 to 3.3 percentage points over progressive prompting. Second, collaborative interaction among three smaller GPT-3.5 models achieved results comparable to, and in some datasets higher than, a single, much larger GPT-4 model. Third, the framework delivered these improvements with lower computational costs—reducing token expenses by 20% compared to a five-sample majority-voting baseline while increasing accuracy by about 3%, and matching the performance of complex ten-path sampling at approximately one-seventh of the cost. Fourth, confidence scoring proved essential, lifting accuracy by an average of 2.92 percentage points by preventing the propagation of erroneous reasoning chains. Finally, performance further improved when mixing distinct model families, such as combining GPT-3.5, GPT-4, and Claude-2.
These findings indicate that collaborative multi-model architectures can reduce dependency on massive, expensive proprietary systems. Organizations can lower operational costs, improve output reliability, and decrease computational overhead by orchestrating smaller or heterogeneous models into structured communication workflows rather than relying solely on larger models or repetitive single-model sampling.
Organizations developing complex reasoning systems should consider adopting multi-model communication paradigms, selecting topologies that align with their operational needs. For example, centralized Report architectures suit setups with one high-capacity model, whereas decentralized Relay or Memory paradigms fit homogeneous systems. Teams should also implement stability-based confidence metrics to filter out early errors and pilot heterogeneous model mixtures to maximize diverse reasoning.
These results are subject to certain boundaries. The experiments evaluated proprietary commercial models and restricted group sizes to three participants to manage token context limits and operational costs. Open-source models and larger model groups were not evaluated. While confidence in the tested commercial benchmarks is high, decision-makers should conduct pilot validations before deploying multi-agent communication frameworks in large-scale production environments.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational Chain-of-Thought prompting methodology upon which Exchange-of-Thought builds its collaborative, cross-model reasoning framework.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2023). Establishes multi-agent debate protocols to enhance reasoning and factuality, directly providing the conceptual basis for collaborative multi-model interaction paradigms.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Formulates reasoning as network graph structures with aggregation and feedback loops, laying the topological foundations extended by Exchange-of-Thought's communication modes.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Extends sequential prompt chains into branching tree structures, motivating the need for structured exploration and multi-path evaluation in complex reasoning.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Demonstrates the power of sampling diverse reasoning paths and voting mechanisms, inspiring confidence evaluation and consensus schemes in multi-agent settings.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Shows how step-by-step reasoning can be elicited zero-shot, establishing the baseline single-model capability that multi-model communication seeks to augment.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Presents structured cross-examination between language models to identify reasoning flaws, directly informing the debate and verification mechanics in multi-model frameworks.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). Explores how multi-agent debate architectures can specifically counter cognitive fixation and Degeneration-of-Thought across complex tasks.
- Paper: Mixture-of-Agents Enhances Large Language Model Capabilities, Junlin Wang et al. (2024). Scales multi-model collaboration into a layered Mixture-of-Agents architecture that aggregates intermediate model representations across diverse open-source LLMs.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). Analyzes how single reasoning models internalize conversational multi-agent dialogue structures into emergent internal 'societies of thought'.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). Extends collaborative multi-LLM problem solving by training a centralized Conductor model to orchestrate agents and dynamic information sharing.
- Paper: LLM Augmented LLMs: Expanding Capabilities through Composition, Rachit Bansal et al. (2024). Develops representation-level cross-attention composition between multiple foundation models to enable capability sharing beyond textual communication.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). Combines multiple reasoning trees into a unified Forest-of-Thought framework to scale test-time compute through collective decision-making.
