Branch-Solve-Merge Improves Large Language Model Evaluation and Generation

Swarnadeep SahaOmer LevyAsli CelikyilmazMohit BansalJason WestonXian Li

article2024NAACL111 citations

Proposes Branch-Solve-Merge, a modular framework that decomposes complex tasks into parallel sub-problems to significantly boost LLM evaluation consistency, mitigate position and length biases, and improve constrained text generation.

Listen

Large language models are increasingly used to evaluate other artificial intelligence outputs and generate complex content. However, these models frequently struggle with multifaceted tasks that require strict adherence to constraints or the balance of diverse criteria, leading to low coherence, poor planning, and evaluation biases. Proprietary models that perform these tasks well can be expensive, while open-source alternatives often lack consistency and reliability. Addressing these issues is vital for organizations seeking cost-effective, dependable, and automated evaluation and generation workflows.

The article introduces and evaluates Branch-Solve-Merge (BSM), a structured program designed to break complex language tasks into independent, parallel sub-tasks. The primary objective is to demonstrate that this modular approach significantly enhances large language model accuracy, reduces systematic biases in evaluation tasks, and improves quality in constrained text generation.

To establish credibility across different model architectures, the researchers evaluated BSM against established baselines using open-source and proprietary models, including Vicuna-33B, LLaMA-2 variants, and GPT-4. The framework decomposes a problem into three modular steps: branching the task into parallel components, solving each sub-problem independently, and merging the intermediate outputs into a final solution. Evaluation performance was tested across eight diverse domains using the MT-Bench dataset, and constrained generation capabilities were assessed using an expanded multi-concept story generation benchmark.

The results show that BSM delivers substantial, measurable improvements across several key dimensions. First, BSM improved agreement between model evaluations and human judgments by up to 26% on conversational benchmarks, enabling the open-source LLaMA-2-70B model to match or exceed the performance of GPT-4 across multiple domains. Second, the framework reduced critical evaluation flaws, cutting position bias by up to 50% and notably decreasing length bias. Third, when applied to GPT-4 evaluating its own outputs, BSM reduced self-enhancement bias and increased human alignment by 3%. Finally, in constrained generation tasks, BSM improved constraint satisfaction by 12% and produced stories that automated judges preferred 93% of the time over baseline generations.

These findings suggest that structured decomposition allows organizations to deploy smaller or open-source models for complex evaluation pipelines, reducing reliance on expensive proprietary interfaces without sacrificing quality. Because the branching module automatically devises task-specific criteria, the framework eliminates the overhead of manually designing evaluation protocols for different subject domains.

Organizations aiming to implement automated assessment or structured generation pipelines should consider adopting modular branching frameworks to improve reliability and lower operating expenses. For high-stakes evaluation environments, combining BSM with multi-sample consistency checks can further minimize position-related errors. Future development should focus on extending the framework into recursive multi-level branching and establishing specialized criteria for safety, toxicity, and bias analysis.

While the reported results demonstrate clear improvements, decision-makers should note certain limitations. The framework requires additional computation due to multiple parallel model invocations, and measuring subjective factors such as length bias remains partly dependent on interpretability assumptions. Nevertheless, the evidence provides strong confidence that parallel task decomposition substantially enhances large language model performance and consistency across varied domains.

arXiv: 2310.15123
Cover for Branch-Solve-Merge Improves Large Language Model Evaluation and Generation

Abstract

Large Language Models (LLMs) are frequently used for multi-faceted language generation and evaluation tasks that involve satisfying intricate user constraints or taking into account multiple aspects and criteria. However, their performance can fall short, due to the model’s lack of coherence and inability to plan and decompose the problem. We propose BRANCH-SOLVE-MERGE (BSM), a Large Language Model program (Schlag et al., 2023) for tackling such challenging natural language tasks. It consists of branch, solve, and merge modules that are parameterized with specific prompts to the base LLM. These three modules plan a decomposition of the task into multiple parallel sub-tasks, independently solve them, and fuse the solutions to the sub-tasks. We apply our method to the tasks of LLM response evaluation and constrained text generation and evaluate its effectiveness with multiple LLMs, including Vicuna, LLaMA-2-chat, and GPT-4. BSM improves the evaluation correctness and consistency for each LLM by enhancing human-LLM agreement by up to 26%, reducing length and pairwise position biases by up to 50%, and allowing LLaMA-2-chat to match or outperform GPT-4 on most domains. On a constraint story generation task, BSM improves the coherence of stories while also improving constraint satisfaction by 12%.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 BRANCH-SOLVE-MERGE
  • 3.1 Components of BRANCH-SOLVE-MERGE
  • 3.2 BSM: Case Study with LLM Evaluation
  • 3.3 BSM: Case Study with Constrained Gen
  • 4 Experiments
  • 4.1 Large Language Model Evaluation
  • 4.1.1 Experimental Setup
  • 4.1.2 Main Results
  • 4.2 Constrained Text Generation
  • 4.2.1 Experimental Setup
  • 4.2.2 Results and Analysis
  • 5 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Additional Experiments: LLM Evaluation
  • A.1 Experimental Setup
  • A.2 Results and Analysis
  • B Additional Experiments: Constrained Text Generation

Knowls

  1. Knowl 1 — Branch-Solve-Merge (BSM) Framework

    model/method

    Branch-Solve-Merge (BSM) is a Large Language Model (LLM) program designed to solve multi-faceted natural language generation and evaluation tasks by decomposing them into parallel sub-tasks. Let pθp_\theta denote an autoregressive language model parameterized by θ\theta, where the probability of generating a sequence of nn tokens x=(x1,…,xn)x = (x_1, \dots, x_n) is:

    pθ(x)=∏i=1npθ(xi∣x1,…,xi−1)p_\theta(x) = \prod_{i=1}^n p_\theta(x_i \mid x_1, \dots, x_{i-1})

    BSM coordinates three modular components via an algorithmic controller Prog:(x,branch(⋅),solve(⋅),merge(⋅))→y\text{Prog}: (x, \text{branch}(\cdot), \text{solve}(\cdot), \text{merge}(\cdot)) \to y that takes an input task instance xx and returns the final output yy:

    1. Branch Module: Decomposes the task input xx into a set of kk independent sub-problems X={x(1),x(2),…,x(k)}∼pθ(X∣promptbranch(x))X = \{x^{(1)}, x^{(2)}, \dots, x^{(k)}\} \sim p_\theta(X \mid \text{prompt}_{\text{branch}}(x)). The model dynamically determines the branching factor kk and the specific sub-tasks based on the input prompt.
    2. Solve Module: Solves each generated sub-problem x(i)x^{(i)} independently and in parallel to produce sub-solution y(i)∼pθ(y(i)∣promptsolve(x(i)))y^{(i)} \sim p_\theta(y^{(i)} \mid \text{prompt}_{\text{solve}}(x^{(i)})) for i∈{1,…,k}i \in \{1, \dots, k\}.
    3. Merge Module: Aggregates the set of intermediate sub-solutions Y={y(1),y(2),…,y(k)}Y = \{y^{(1)}, y^{(2)}, \dots, y^{(k)}\} to produce a unified global output y∼pθ(y∣promptmerge(Y))y \sim p_\theta(y \mid \text{prompt}_{\text{merge}}(Y)). The merge step can be implemented using deterministic non-neural aggregation (e.g., score summation) or neural LLM aggregation that reasons across or fuses sub-solutions.
  2. Knowl 2 — BSM Algorithm for Pairwise LLM Response Evaluation

    algorithm

    The BSM algorithm for pairwise LLM response evaluation determines which of two candidate LLM responses, r(A)r^{(A)} or r(B)r^{(B)}, is superior for a given multi-turn user question qq (comprising turn 1 question q1q_1 and optional turn 2 follow-up q2q_2).

    To eliminate position bias, the algorithm executes two evaluation passes: forward order (r(A),r(B))(r^{(A)}, r^{(B)}) and reverse order (r(B),r(A))(r^{(B)}, r^{(A)}). In each pass, the branch module conditions on qq to generate kk question-specific criteria C={c1,…,ck}C = \{c_1, \dots, c_k\} (such as relevance, clarity, or code correctness). For each criterion cic_i, the solve module assigns individual scores (si(1),si(2))∈[1,5]2(s_i^{(1)}, s_i^{(2)}) \in [1, 5]^2 alongside qualitative rationales. The merge module sums the scores across all criteria for each candidate response: S(1)=∑i=1ksi(1)S^{(1)} = \sum_{i=1}^k s_i^{(1)} and S(2)=∑i=1ksi(2)S^{(2)} = \sum_{i=1}^k s_i^{(2)}. If S(1)>S(2)S^{(1)} > S^{(2)}, the first response wins; if S(2)>S(1)S^{(2)} > S^{(1)}, the second response wins; otherwise, it is a tie. The controller assigns a final verdict of AA or BB if and only if both passes agree; otherwise, a tie is declared.

    Input: Question qq, Candidate responses r(A)r^{(A)} and r(B)r^{(B)}, Modules $m = (\text{branch}, \text{solve}, \text{merge})
    Output: Final preference judgment $y \in \{A, B, \text{tie}\}
    function EvaluateBSM(qq, r1r_1, r2r_2, mm)
        C←branch(q)C \leftarrow \text{branch}(q)
        for each ci∈Cc_i \in C do
            (si(1),si(2))←solve(q,r1,r2,ci)(s_i^{(1)}, s_i^{(2)}) \leftarrow \text{solve}(q, r_1, r_2, c_i)
        end for
        yrun←merge(q,C,{si(1)}i=1∣C∣,{si(2)}i=1∣C∣)y_{\text{run}} \leftarrow \text{merge}(q, C, \{s_i^{(1)}\}_{i=1}^{|C|}, \{s_i^{(2)}\}_{i=1}^{|C|})
        return yruny_{\text{run}}
    end function
    y(A,B)←EvaluateBSM(q,r(A),r(B),m)y^{(A, B)} \leftarrow \text{EvaluateBSM}(q, r^{(A)}, r^{(B)}, m)
    y(B,A)←EvaluateBSM(q,r(B),r(A),m)y^{(B, A)} \leftarrow \text{EvaluateBSM}(q, r^{(B)}, r^{(A)}, m)
    if y(A,B)=firsty^{(A, B)} = \text{first} and y(B,A)=secondy^{(B, A)} = \text{second} then
        y←Ay \leftarrow A
    else if y(A,B)=secondy^{(A, B)} = \text{second} and y(B,A)=firsty^{(B, A)} = \text{first} then
        y←By \leftarrow B
    else
        y←tiey \leftarrow \text{tie}
    end if
    return yy
  3. Knowl 3 — BSM for Constrained Story Generation

    model/method

    BSM addresses constrained text generation by decomposing a large constraint set into smaller subsets and maintaining narrative coherence through global topic planning. Given a target concept set l={w1,w2,…,wm}l = \{w_1, w_2, \dots, w_m\}, the method operates through three modules:

    1. Branch Module: branch(l)→(l1,l2,t)\text{branch}(l) \to (l_1, l_2, t) partitions the concept set into two disjoint subsets l1l_1 and l2l_2 such that l1∪l2=ll_1 \cup l_2 = l and l1∩l2=∅l_1 \cap l_2 = \emptyset, and simultaneously generates an overarching narrative topic tt.
    2. Solve Module: solve(li,t)→yi\text{solve}(l_i, t) \to y_i generates an intermediate single-paragraph story yiy_i that satisfies all concepts in subset lil_i while adhering to topic tt (i∈{1,2}i \in \{1, 2\}). Solving on a smaller subset simplifies constraint satisfaction for the LLM.
    3. Merge Module: merge(y1,y2,l1,l2)→y\text{merge}(y_1, y_2, l_1, l_2) \to y fuses the two intermediate stories y1y_1 and y2y_2 into a single coherent paragraph yy while ensuring that all concepts from both l1l_1 and l2l_2 are preserved. Because both sub-stories share the same topic tt, semantic and stylistic continuity is maintained in the final story.
  4. Knowl 4 — Evaluation Metrics for LLM Evaluators

    definition

    Pairwise evaluation performance and evaluator reliability are measured using four complementary metrics:

    • LLM-Human Agreement (Ag∈[0,1]Ag \in [0, 1]): The fraction of evaluation instances where the LLM evaluator's preference verdict matches the human expert consensus.
    • Position Bias (PB∈[0,100]%PB \in [0, 100]\%): The fraction of evaluation instances where swapping the presentation order of the two candidate responses changes the evaluator's preferred verdict (lower values indicate higher consistency).
    • Length Bias (LB∈[0,100]%LB \in [0, 100]\%): The percentage of instances where human annotators prefer the shorter response, but the LLM evaluator chooses the longer response (lower values indicate less distortion toward verbosity).
    • Self-Enhancement Bias (SBSB): The tendency of an LLM evaluator to favor its own outputs. It is assessed by evaluating LLM-human agreement on the subset of test samples where one of the candidate responses was generated by the evaluator model itself.
  5. Knowl 5 — Pairwise Evaluation Performance of BSM vs. Baselines on MT-Bench Writing

    data/table

    The performance of BSM with LLaMA-2-70B-chat as the base model was evaluated on the 300 pairwise comparison instances of the 'Writing' category from the MT-Bench dataset. BSM was compared against zero-shot prompting (relative preference and absolute scoring), Plan&Solve prompting, and Self-Consistency (sampling matching the branching factor of BSM with majority voting).

    Method Overall Turn-1 Turn-2
    Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow
    Zero-shot (Relative) 0.43 51.66 54.88 0.53 42.66 50.00 0.34 60.66 59.42
    Zero-shot (Absolute) 0.45 30.00 48.87 0.56 19.33 43.75 0.34 40.66 53.62
    PlanSolve 0.43 43.00 54.13 0.43 42.00 51.56 0.43 44.00 56.52
    Self-Consistency 0.52 35.66 48.12 0.57 32.00 45.31 0.47 39.33 50.72
    BSM 0.55 17.33 39.09 0.60 14.66 39.46 0.50 20.00 39.13

    BSM outperforms all baselines in human agreement (Ag=0.55Ag=0.55 overall) while substantially reducing position bias (PB=17.33%PB=17.33\% vs 51.66%51.66\% for zero-shot relative) and length bias (LB=39.09%LB=39.09\% vs 54.88%54.88\%). On turn-2 questions, where baseline agreement degrades sharply (from 0.530.53 to 0.340.34 for zero-shot relative), BSM maintains Ag=0.50Ag=0.50, indicating that dynamic criterion planning is especially beneficial in long-context multi-turn evaluation.

  6. Knowl 6 — Evaluation Performance Across Base LLMs and Self-Enhancement Bias Reduction

    data/table

    BSM was evaluated across four base LLMs of varying size on MT-Bench writing questions (300 pairs) to measure generalizability and bias mitigation.

    Method Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow
    Zero-shot (w/ LLaMA-2-7B-chat) 0.39 62.33 54.88
    BSM (w/ LLaMA-2-7B-chat) 0.41 48.33 53.38
    Zero-shot (w/ Vicuna-33B) 0.51 30.66 48.12
    BSM (w/ Vicuna-33B) 0.56 20.00 42.85
    Zero-shot (w/ LLaMA-2-70B-chat) 0.43 51.66 54.88
    BSM (w/ LLaMA-2-70B-chat) 0.55 17.33 39.09
    Zero-shot (w/ GPT-4) 0.59 17.33 39.09
    BSM (w/ GPT-4) 0.62 17.00 36.84

    BSM improves human agreement (AgAg) across all base models and reduces PBPB and LBLB across open-source models. Applying BSM to LLaMA-2-70B-chat increases agreement by 12%12\% absolute, approaching GPT-4 zero-shot performance.

    In self-enhancement bias experiments on MT-Bench subsets where one candidate response is generated by GPT-4 and judged by GPT-4:

    • Zero-shot GPT-4 achieves Ag=0.51Ag = 0.51, PB=6.33%PB = 6.33\%, LB=36.36%LB = 36.36\%.
    • BSM with GPT-4 achieves Ag=0.54Ag = 0.54, PB=7.33%PB = 7.33\%, LB=34.54%LB = 34.54\%. This demonstrates that BSM improves alignment with human preferences even when an LLM evaluates its own outputs.
  7. Knowl 7 — BSM Performance on Reference-Based and Multi-Domain MT-Bench Categories

    data/table

    BSM with LLaMA-2-70B-chat was evaluated across reference-based categories (Coding, Reasoning, Math; conditioned on a curated reference answer) and multi-domain categories (Roleplay, Extraction, STEM, Humanities) on MT-Bench.

    Category Method Overall Turn-1 Turn-2
    Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow Ag ↑\uparrow PB ↓\downarrow LB ↓\downarrow
    Coding Zero-shot (LLaMA-2-70B) 0.47 52.33 51.32 - - - - - -
    BSM (LLaMA-2-70B) 0.61 25.66 42.47 - - - - - -
    GPT-4 0.61 19.66 38.93 - - - - - -
    Reasoning Zero-shot (LLaMA-2-70B) 0.47 38.00 48.75 - - - - - -
    BSM (LLaMA-2-70B) 0.57 20.33 46.25 - - - - - -
    GPT-4 0.64 22.66 53.75 - - - - - -
    Math Zero-shot (LLaMA-2-70B) 0.52 45.66 50.56 - - - - - -
    BSM (LLaMA-2-70B) 0.64 17.66 34.83 - - - - - -
    GPT-4 0.62 19.00 39.32 - - - - - -
    Roleplay Zero-shot (LLaMA-2-70B) 0.55 29.66 51.67 0.61 30.00 48.14 0.50 29.33 55.88
    BSM (LLaMA-2-70B) 0.61 11.00 40.26 0.66 10.66 38.27 0.56 11.33 42.64
    GPT-4 0.64 13.66 43.62 0.65 16.00 45.67 0.63 11.33 41.17
    Extraction Zero-shot (LLaMA-2-70B) 0.40 70.66 51.82 0.46 61.33 51.47 0.33 80.00 52.08
    BSM (LLaMA-2-70B) 0.55 31.33 40.24 0.55 32.00 45.58 0.44 30.66 36.45
    GPT-4 0.71 15.00 33.53 0.68 13.33 35.29 0.75 16.66 32.29
    STEM Zero-shot (LLaMA-2-70B) 0.46 59.33 55.31 0.50 52.66 51.19 0.43 66.00 61.40
    BSM (LLaMA-2-70B) 0.72 10.33 44.68 0.70 10.66 40.47 0.73 10.00 50.87
    GPT-4 0.72 13.66 46.80 0.68 16.66 44.04 0.75 10.66 50.87
    Humanities Zero-shot (LLaMA-2-70B) 0.46 59.00 45.69 0.51 52.00 49.18 0.41 66.00 43.33
    BSM (LLaMA-2-70B) 0.67 18.00 36.42 0.63 18.00 39.34 0.71 18.00 34.44
    GPT-4 0.73 14.00 37.08 0.70 19.33 42.62 0.76 8.66 33.33

    In Math, BSM with LLaMA-2-70B-chat outperforms GPT-4 on agreement (Ag=0.64Ag=0.64 vs 0.620.62), position bias (PB=17.66%PB=17.66\% vs 19.00%19.00\%), and length bias (LB=34.83%LB=34.83\% vs 39.32%39.32\%). In STEM, BSM achieves a 26%26\% absolute gain over zero-shot (0.46→0.720.46 \to 0.72), matching GPT-4 agreement while achieving lower biases.

  8. Knowl 8 — Constrained Story Generation Results and Error Analysis

    data/table

    Constrained story generation was evaluated on 100 instances from CommonGen-Hard requiring the inclusion of 10 concepts per story. Performance was measured by All Present (AP ↑\uparrow, percentage of stories containing all 10 concepts) and Missing Concepts (MC ↓\downarrow, average percentage of missing concepts).

    Method LLaMA-2-7B-chat LLaMA-2-70B-chat
    AP ↑\uparrow MC ↓\downarrow AP ↑\uparrow MC ↓\downarrow
    Zero-shot 15.0% 17.3% 22.0% 27.2%
    PlanSolve 13.0% 18.0% 21.0% 26.6%
    Self-Consistency 19.0% 16.6% 24.0% 20.1%
    BSM 23.0% 15.5% 28.0% 14.7%

    BSM achieves superior constraint satisfaction over all baselines for both model scales (28.0%28.0\% AP and 14.7%14.7\% MC with LLaMA-2-70B-chat).

    In a pairwise story quality evaluation judged by GPT-4 (accounting for order swap consistency), stories generated by BSM with LLaMA-2-70B-chat are preferred 93%93\% of the time over zero-shot baselines.

    Error analysis of the 72%72\% of BSM stories missing at least one concept revealed:

    • 60%60\% of errors originate in the solve module (concepts omitted during intermediate branch generation).
    • 12%12\% of errors originate in the merge module (concepts present in intermediate sub-stories but lost during fusion).
  9. Knowl 9 — Impact of Branching Factor, Scoring Scale, and Branch-Level Self-Consistency

    empirical result

    Ablations on 100 MT-Bench writing samples evaluated with LLaMA-2-70B-chat show:

    1. Branching Factor (BFBF): Varying the maximum branching factor in the branch prompt yields:
      • BF=2BF=2: Ag=0.50Ag = 0.50, PB=22.00%PB = 22.00\%, LB=49.20%LB = 49.20\%
      • BF=3BF=3: Ag=0.52Ag = 0.52, PB=19.00%PB = 19.00\%, LB=38.09%LB = 38.09\%
      • BF=4BF=4: Ag=0.53Ag = 0.53, PB=19.00%PB = 19.00\%, LB=38.09%LB = 38.09\%
      • BF=5BF=5: Ag=0.52Ag = 0.52, PB=12.00%PB = 12.00\%, LB=34.92%LB = 34.92\% Human agreement peaks at BF=4BF=4 and saturates, whereas position bias monotonically decreases with more branches due to variance reduction.
    2. Evaluation Scale: Varying the score range in the solve prompt between 1–51\text{--}5 and 1–101\text{--}10 yields comparable overall agreement (0.520.52 vs 0.500.50) and length bias (34.92%34.92\% vs 36.50%36.50\%), though 1–101\text{--}10 slightly increases position bias (12.00%→18.00%12.00\% \to 18.00\%).
    3. Branch-Level Self-Consistency (BSM+SC): Sampling 5 evaluations per branch (temperature 0.7) and averaging scores across runs preserves overall agreement (Ag=0.55Ag=0.55) while providing an additional 2%2\% absolute reduction in position bias (PB=17.33%→15.33%PB=17.33\% \to 15.33\%).
  10. Knowl 10 — Stated Limitations of Branch-Solve-Merge

    limitation

    The authors identify four main limitations of the Branch-Solve-Merge framework:

    1. Safety and Toxicity: The study does not evaluate safety, toxicity, or alignment-related biases in LLM outputs.
    2. Isolating Length Bias: Measuring length bias in isolation is challenging because human evaluators naturally favor longer, more detailed responses for open-ended queries, making it difficult to separate spurious bias from genuine quality preferences.
    3. Absence of Recursive/Hierarchical Decomposition: While error analysis shows that sub-task solver failures account for 60%60\% of constraint omissions in generation, recursive multi-level branching was not explored due to the increased computational cost of invoking multiple base LLM calls.
    4. Efficiency vs. Performance Trade-off: Although parallel sub-task decomposition is theoretically amenable to parallel execution, this work focused strictly on task performance rather than computational latency or runtime efficiency.

Coverage note — None was omitted.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022b. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  3. 3.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, et al. 2023. Graph of thoughts: Solving elaborate problems with large language models. arXiv preprint arXiv:2308.09687.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  6. 6.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  7. 7.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2023. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109.
  8. 8.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128.
  9. 9.Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15607–15631, Toronto, Canada. Association for Computational Linguistics.
  10. 10.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  11. 11.Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Visual programming for text-to-image generation and evaluation. In NeurIPS.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  13. 13.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  14. 14.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. arXiv preprint arXiv:2208.14271.
  15. 15.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  16. 16.David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A Saurous, Jascha Sohl-Dickstein, et al. 2022. Language model cascades. arXiv preprint arXiv:2207.10342.
  17. 17.Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. 2022. Successive prompting for decomposing complex questions. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1251–1265.
  18. 18.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback. arXiv preprint arXiv:2305.14387.
  19. 19.Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamile Lukoši ˙ ut¯ e, Anna Chen, Anna ˙ Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459.
  20. 20.Tanmay Gupta and Aniruddha Kembhavi. 2023. Visual programming: Compositional visual reasoning without training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962.
  21. 21.Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2023. Are large language model-based evaluators the solution to scaling up multilingual evaluation? arXiv preprint arXiv:2309.07462.
  22. 22.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798.
  23. 23.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR.
  24. 24.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  25. 25.Tushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2021. Text modular networks: Learning to decompose tasks in the language of existing models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1264–1279.
  26. 26.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  27. 27.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al. 2023. OpenAssistant conversations–democratizing large language model alignment. arXiv preprint arXiv:2304.07327.
  28. 28.Bin Lei, Chunhua Liao, Caiwen Ding, et al. 2023. Boosting logical reasoning in large language models through a new framework: The graph of thought. arXiv preprint arXiv:2308.08614.
  29. 29.Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Luke Zettlemoyer, Omer Levy, Jason Weston, and Mike Lewis. 2023. Self-alignment with instruction backtranslation. arXiv preprint arXiv:2308.06259.
  30. 30.Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. 2022. Diffusionlm improves controllable text generation. Advances in Neural Information Processing Systems, 35:4328–4343.
  31. 31.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  32. 32.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840.
  33. 33.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  34. 34.Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, et al. 2022. Neurologic a* esque decoding: Constrained text generation with lookahead heuristics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 780–799.
  35. 35.Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Neurologic decoding:(un) supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4299.
  36. 36.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651.
  37. 37.Xuefei Ning, Zinan Lin, Zixuan Zhou, Huazhong Yang, and Yu Wang. 2023. Skeleton-of-thought: Large language models can do parallel decoding. arXiv preprint arXiv:2307.15337.
  38. 38.OpenAI. 2023a. Evals is a framework for evaluating llms and llm systems, and an open-source registry of benchmarks.
  39. 39.OpenAI. 2023b. Gpt-4 technical report.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  42. 42.Swarnadeep Saha, Xinyan Yu, Mohit Bansal, Ramakanth Pasunuru, and Asli Celikyilmaz. 2023. MURMUR: Modular multi-step reasoning for semistructured data-to-text generation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 11069–11090, Toronto, Canada. Association for Computational Linguistics.
  43. 43.Swarnadeep Saha, Shiyue Zhang, Peter Hase, and Mohit Bansal. 2022. Summarization programs: Interpretable abstractive summarization with neural modular trees. In The Eleventh International Conference on Learning Representations.
  44. 44.Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz, Wen-tau Yih, Jason Weston, Jürgen Schmidhuber, and Xian Li. 2023. Large language model programs. arXiv preprint arXiv:2305.05364.
  45. 45.Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022. Human evaluation of conversations is an open problem: comparing the sensitivity of various methods for evaluating dialogue agents. In Proceedings of the 4th Workshop on NLP for Conversational AI, pages 77–97.
  46. 46.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  47. 47.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023a. Plan-and-solve prompting: Improving zeroshot chain-of-thought reasoning by large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2609–2634, Toronto, Canada. Association for Computational Linguistics.
  48. 48.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926.
  49. 49.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  50. 50.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al. 2023c. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv:2306.04751.
  51. 51.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  52. 52.Minghao Wu and Alham Fikri Aji. 2023. Style over substance: Evaluation biases for large language models. arXiv preprint arXiv:2307.03025.
  53. 53.Shunyu Yao, Howard Chen, Austin W Hanjie, Runzhe Yang, and Karthik Narasimhan. 2023a. COLLIE: Systematic construction of constrained text generation tasks. arXiv preprint arXiv:2307.08689.
  54. 54.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023b. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  55. 55.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
  56. 56.Yao Yao, Zuchao Li, and Hai Zhao. 2023c. Beyond chain-of-thought, effective graph-of-thought reasoning in large language models. arXiv preprint arXiv:2305.16582.
  57. 57.Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023. Wider and deeper llm networks are fairer llm evaluators. arXiv preprint arXiv:2308.01862.
  58. 58.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  59. 59.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  60. 60.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  61. 61.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, et al. 2022. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Saha, S., et al. “Branch-Solve-Merge Improves Large Language Model Evaluation and Generation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 8352–70, https://doi.org/10.18653/v1/2024.naacl-long.462.
APA
Saha, S., Levy, O., Celikyilmaz, A., Bansal, M., Weston, J., & Li, X. (2024). Branch-Solve-Merge Improves Large Language Model Evaluation and Generation. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8352–8370. https://doi.org/10.18653/v1/2024.naacl-long.462
Chicago
Saha, S., O. Levy, A. Celikyilmaz, M. Bansal, J. Weston, and X. Li. 2024. “Branch-Solve-Merge Improves Large Language Model Evaluation and Generation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8352–70. https://doi.org/10.18653/v1/2024.naacl-long.462.
Harvard
Saha, S. et al. (2024) “Branch-Solve-Merge Improves Large Language Model Evaluation and Generation”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8352–8370. Available at: https://doi.org/10.18653/v1/2024.naacl-long.462.
Vancouver
1. Saha S, Levy O, Celikyilmaz A, Bansal M, Weston J, Li X (2024) Branch-Solve-Merge Improves Large Language Model Evaluation and Generation. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 8352–8370

BibTeX

@inproceedings{saha-etal-2024-branch,
    title = "Branch-Solve-Merge Improves Large Language Model Evaluation and Generation",
    author = "Saha, Swarnadeep  and
      Levy, Omer  and
      Celikyilmaz, Asli  and
      Bansal, Mohit  and
      Weston, Jason  and
      Li, Xian",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.462/",
    doi = "10.18653/v1/2024.naacl-long.462",
    pages = "8352--8370"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/