Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
Swarnadeep SahaOmer LevyAsli CelikyilmazMohit BansalJason WestonXian Li
Proposes Branch-Solve-Merge, a modular framework that decomposes complex tasks into parallel sub-problems to significantly boost LLM evaluation consistency, mitigate position and length biases, and improve constrained text generation.
Large language models are increasingly used to evaluate other artificial intelligence outputs and generate complex content. However, these models frequently struggle with multifaceted tasks that require strict adherence to constraints or the balance of diverse criteria, leading to low coherence, poor planning, and evaluation biases. Proprietary models that perform these tasks well can be expensive, while open-source alternatives often lack consistency and reliability. Addressing these issues is vital for organizations seeking cost-effective, dependable, and automated evaluation and generation workflows.
The article introduces and evaluates Branch-Solve-Merge (BSM), a structured program designed to break complex language tasks into independent, parallel sub-tasks. The primary objective is to demonstrate that this modular approach significantly enhances large language model accuracy, reduces systematic biases in evaluation tasks, and improves quality in constrained text generation.
To establish credibility across different model architectures, the researchers evaluated BSM against established baselines using open-source and proprietary models, including Vicuna-33B, LLaMA-2 variants, and GPT-4. The framework decomposes a problem into three modular steps: branching the task into parallel components, solving each sub-problem independently, and merging the intermediate outputs into a final solution. Evaluation performance was tested across eight diverse domains using the MT-Bench dataset, and constrained generation capabilities were assessed using an expanded multi-concept story generation benchmark.
The results show that BSM delivers substantial, measurable improvements across several key dimensions. First, BSM improved agreement between model evaluations and human judgments by up to 26% on conversational benchmarks, enabling the open-source LLaMA-2-70B model to match or exceed the performance of GPT-4 across multiple domains. Second, the framework reduced critical evaluation flaws, cutting position bias by up to 50% and notably decreasing length bias. Third, when applied to GPT-4 evaluating its own outputs, BSM reduced self-enhancement bias and increased human alignment by 3%. Finally, in constrained generation tasks, BSM improved constraint satisfaction by 12% and produced stories that automated judges preferred 93% of the time over baseline generations.
These findings suggest that structured decomposition allows organizations to deploy smaller or open-source models for complex evaluation pipelines, reducing reliance on expensive proprietary interfaces without sacrificing quality. Because the branching module automatically devises task-specific criteria, the framework eliminates the overhead of manually designing evaluation protocols for different subject domains.
Organizations aiming to implement automated assessment or structured generation pipelines should consider adopting modular branching frameworks to improve reliability and lower operating expenses. For high-stakes evaluation environments, combining BSM with multi-sample consistency checks can further minimize position-related errors. Future development should focus on extending the framework into recursive multi-level branching and establishing specialized criteria for safety, toxicity, and bias analysis.
While the reported results demonstrate clear improvements, decision-makers should note certain limitations. The framework requires additional computation due to multiple parallel model invocations, and measuring subjective factors such as length bias remains partly dependent on interpretability assumptions. Nevertheless, the evidence provides strong confidence that parallel task decomposition substantially enhances large language model performance and consistency across varied domains.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This paper establishes the LLM-as-a-judge paradigm and the MT-Bench benchmark used directly by the source to evaluate and mitigate model evaluation biases.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). This work introduces graph-structured task decomposition and thought aggregation, providing the conceptual foundation for branching, solving, and merging sub-tasks in LLM generation and reasoning.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). This foundational paper presents self-consistency via multi-path sampling, a core concept underlying parallel reasoning aggregation and consensus mechanisms in modular prompting.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2023). This paper demonstrates how decomposing problem-solving across multiple collaborating agents mitigates single-model reasoning biases and improves factual accuracy.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). This work formulates criterion-based and chain-of-thought evaluation protocols with large language models that Branch-Solve-Merge modularizes and automates.
- Paper: Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness, Jiuhai Chen et al. (2024). This paper introduces methods for quantifying uncertainty and evaluating self-consistency in black-box LLMs, which the source relies on for multi-sample validation and bias mitigation.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This comprehensive survey synthesizes advanced judging architectures, position-bias corrections, and multi-agent setups that extend modular evaluation strategies like Branch-Solve-Merge.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). This paper introduces instance-specific fine-grained rubrics for LLM evaluation, providing a broader benchmark framework to apply and test modular, multi-criteria evaluation pipelines.
- Paper: Mixture-of-Agents Enhances Large Language Model Capabilities, Junlin Wang et al. (2024). This work extends multi-agent decomposition and synthesis into layered Mixture-of-Agents architectures for collaborative LLM generation and evaluation.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This study evaluates LLM judges on complex constraint adherence using adversarial benchmarks, offering a direct evaluation testbed for structured decomposition frameworks.
- Paper: Self-Generated Critiques Boost Reward Modeling for Language Models, Yue Yu 0009 et al. (2025). This research builds on structured critique generation to improve reward modeling and preference alignment without relying on proprietary teacher models.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). This paper scales test-time compute across multiple reasoning trees and consensus-guided merging, generalizing branching-and-merging mechanisms to hard reasoning tasks.
- Paper: Self-Improving Language Models with Bidirectional Evolutionary Search, Guowei Xu et al. (2026). This work develops bidirectional evolutionary search to decompose tasks backward and merge trajectory segments forward, advancing modular generation and refinement.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). This work presents an automated framework to divide, assign, and orchestrate subtasks across diverse LLMs, automating the decomposition and merging process.
