LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Dongfu JiangXiang RenBill Yuchen Lin

article2023ACL729 citations

Proposes LLM-Blender, an ensembling framework that combines multiple open-source language models by using cross-attention pairwise ranking to select top candidate outputs and a generative fusion module to merge them into superior responses.

Listen

Organizations deploying artificial intelligence increasingly rely on open-source large language models. However, individual language models exhibit distinct strengths and weaknesses due to variations in architectures, training datasets, and fine-tuning procedures. As a result, no single model consistently delivers the best answer across all user prompts, leading to suboptimal quality, increased error rates, and heightened exposure to model biases when relying on a single fixed system.

The article demonstrates and evaluates an ensemble framework, termed LLM-BLENDER, designed to combine outputs from multiple open-source language models dynamically. The primary objective is to reliably generate superior responses by systematically ranking candidate outputs and fusing the best candidates into a single refined answer.

To establish a standard evaluation, the researchers introduced MixInstruct, a benchmark consisting of 110,000 instruction-following examples evaluated across 11 popular open-source models, including Vicuna, OpenAssistant, and Alpaca. The proposed framework implements a two-stage pipeline: first, a ranking module directly compares candidate pairs alongside the source input to assess subtle quality differences and rank them; second, a generative fusion module combines the top three ranked outputs to synthesize a unified response. The approach was evaluated against standalone language models and traditional individual scoring techniques using both automated linguistic metrics and automated comparative rankings.

The findings confirm that model performance varies significantly by query; the top-performing individual model (Vicuna) achieved the best answer in only 21.22% of examples. The proposed pairwise ranker significantly outperformed standard reranking methods, achieving an average rank of 3.20 compared to the top standalone model's 3.90—an approximate 18% relative improvement. When integrating both ranking and generative fusion, the full framework achieved an average rank of 3.01 and placed within the top three outputs across 68.59% of evaluated cases, compared to 52.88% for the best individual model. Furthermore, generative fusion consistently exceeded individual models across all standard automated text-quality benchmarks.

These results indicate that ensembling open-source models mitigates individual system failures, elevates output consistency, and reduces reliance on proprietary closed-source solutions. Rather than searching for a single universal model, organizations can achieve higher reliability and alignment with human expectations by combining smaller, specialized models. However, generating outputs across multiple models and performing pairwise comparisons increases computational overhead and latency, introducing practical trade-offs for real-time operations.

Technical leaders and product teams should consider deploying dynamic ensemble workflows where response accuracy and safety outweigh strict latency constraints. For latency-sensitive deployments, the article supports adopting accelerated comparison methods, such as single-pass sorting algorithms or parallelized pairwise evaluations, to balance performance with compute efficiency. Further testing in organization-specific domains and live user pilots is recommended before full-scale deployment.

Readers should note two main limitations: the computational requirement scales quadratically with candidate counts during full pairwise evaluation, and large-scale validation relied primarily on automated and language-model-based judging rather than extensive human evaluation. While automated evaluations demonstrate strong alignment with comparative benchmarks, results should be interpreted with measured caution until verified against target-domain human review.

Cover for LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Abstract

We present LLM-BLENDER, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs). Our framework consists of two modules: PAIRRANKER and GENFUSER, addressing the observation that optimal LLMs for different examples can significantly vary. PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs. It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one. Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking. Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses. To facilitate large-scale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons. Our LLM-BLENDER significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Problem Setup
  • 2.2 MixInstruct: A New Benchmark
  • 2.3 LLM-BLENDER: A Novel Framework
  • 3 PAIRRANKER: Pairwise Ranking
  • 3.1 Baseline Methods
  • 3.2 Pairwise Comparisons
  • 3.3 PAIRRANKER Architecture
  • 4 GENFUSER: Generative Fusion
  • 5 Evaluation
  • 5.1 Setup
  • 5.2 Main results
  • 6 Related Work
  • 7 Conclusion & Future Directions
  • Limitations
  • Ethical Statement
  • Acknowledgements
  • References
  • ACL 2023 Responsible NLP Checklist
  • A For every submission:
  • B Did you use or create scientific artifacts?
  • C Did you run computational experiments?
  • D Did you use human annotators (e.g., crowdworkers) or research with human participants?

Knowls

  1. Knowl 1 — LLM-BLENDER Framework for Ensembling Large Language Models

    model/method

    LLM-BLENDER is an ensembling framework designed to combine the strengths of multiple open-source large language models (LLMs) for instruction-following tasks. It operates in two sequential stages:

    1. Pairwise Ranking (PAIRRANKER): Given an input instruction xx and a candidate output set Y={y1,…,yN}Y = \{y_1, \dots, y_N\} generated by NN different LLMs, PAIRRANKER jointly encodes xx and pairs of candidates using a cross-attention transformer to capture fine-grained distinctions and rank the candidates.

    2. Generative Fusion (GENFUSER): Given the top KK ranked candidates (K<NK < N, typically K=3K=3) selected by PAIRRANKER and the input xx, a sequence-to-sequence model merges them into an improved final output response y^\hat{y} that capitalizes on candidate strengths while mitigating individual errors.

  2. Knowl 2 — MixInstruct Benchmark Dataset

    experimental setup

    MixInstruct is an instruction-following benchmark dataset designed to train and evaluate ensemble methods for LLMs.

    • Data Sources and Composition: It comprises 110,000 instruction-following examples sampled from four primary sources: Alpaca-GPT4 (22,862 examples; source GPT-4; average 22 input and 48 output tokens), Dolly-15K (7,584 examples; source Human; 24 input and 53 output tokens), GPT4All-LAION (76,552 examples; source ChatGPT; 18 input and 72 output tokens), and ShareGPT (3,002 examples; source ChatGPT; 36 input and 63 output tokens). The overall dataset averages 20 input and 66 output tokens per instance.

    • Dataset Splits: Randomly split into 100,000 training instances, 5,000 validation instances, and 5,000 test instances.

    • Candidate Models: Outputs are generated across all 110k examples using N=11N=11 open-source LLMs: OpenAssistant, Vicuna, Alpaca, Baize, MOSS, ChatGLM, Koala, Dolly V2, Mosaic MPT, StableLM, and Flan-T5.

    • Oracle Evaluation: For test instances, oracle pairwise evaluations across all 11×10/2=5511 \times 10 / 2 = 55 candidate pairs per input are generated using comparative prompts with ChatGPT (GPT-Rank).

  3. Knowl 3 — PAIRRANKER Cross-Attention Architecture and Pairwise Encoding

    model/method

    PAIRRANKER compares candidate responses in pairs (yi,yj)(y_i, y_j) conditioned on the user prompt xx using a shared cross-attention encoder (such as DeBERTa-v3 with 400M parameters).

    • Input Representation: The prompt and two candidate outputs are concatenated into a single sequence using special delimiter tokens: "<s><source> x </s> <candidate1> yi </s> <candidate2> yj </s>"\text{"<s><source> } x \text{ </s> <candidate1> } y_i \text{ </s> <candidate2> } y_j \text{ </s>"}

    • Token Embeddings: The encoder output embeddings corresponding to the separator tokens ⟨source⟩\langle\text{source}\rangle, ⟨candidate1⟩\langle\text{candidate1}\rangle, and ⟨candidate2⟩\langle\text{candidate2}\rangle represent the vector embeddings x\mathbf{x}, yi\mathbf{y}_i, and yj\mathbf{y}_j, respectively.

    • Scoring Layer: The concatenated vectors [x;yi][\mathbf{x}; \mathbf{y}_i] and [x;yj][\mathbf{x}; \mathbf{y}_j] are each passed through a shared multi-layer perceptron (MLP) head with output dimension equal to the number of target quality functions QQ. The final pair-specific scores s(i,j)is_{(i,j)}^i and s(i,j)js_{(i,j)}^j are computed by averaging the predicted scores across all QQ.

    • Pairwise Superiority: The model's confidence that yiy_i is better than yjy_j is given by sij=s(i,j)i−s(i,j)js_{ij} = s_{(i,j)}^i - s_{(i,j)}^j.

  4. Knowl 4 — PAIRRANKER Multi-Objective Pairwise Ranking Loss

    equation

    To train PAIRRANKER to distinguish candidate quality based on reference quality functions QQ (such as BARTScore or BERTScore against ground truth yy), the model is optimized via multi-task binary cross-entropy classification.

    For an input xx, reference yy, candidate pair (yi,yj)(y_i, y_j), and quality metric QQ, the loss is: LQ=−zilog⁡σ(s(i,j)i)−(1−zj)log⁡σ(s(i,j)j)L_Q = -z_i \log \sigma(s_{(i,j)}^i) - (1 - z_j) \log \sigma(s_{(i,j)}^j) where σ(v)=11+e−v\sigma(v) = \frac{1}{1 + e^{-v}} is the sigmoid function, s(i,j)is_{(i,j)}^i and s(i,j)js_{(i,j)}^j are predicted candidate scores, and the binary targets (zi,zj)(z_i, z_j) are defined as: (zi,zj)={(1,0),if Q(yi,y)≥Q(yj,y)(0,1),if Q(yi,y)<Q(yj,y)(z_i, z_j) = \begin{cases} (1, 0), & \text{if } Q(y_i, y) \ge Q(y_j, y) \\ (0, 1), & \text{if } Q(y_i, y) < Q(y_j, y) \end{cases} When optimizing across multiple quality functions QQ, the total loss is the average over all metrics: L=∑QLQ\mathcal{L} = \sum_Q L_Q

  5. Knowl 5 — Candidate Ranking Aggregation Algorithms in PAIRRANKER

    algorithm

    From all candidate pairs in YY, PAIRRANKER constructs a pairwise comparison matrix M∈RN×N\mathbf{M} \in \mathbb{R}^{N \times N}, where Mij=sijM_i^j = s_{ij} denotes the confidence that candidate yiy_i is superior to yjy_j. The final candidate ordering is computed using one of three aggregation methods:

    1. MaxLogits (O(N2)\mathcal{O}(N^2) comparisons): Evaluates the sum of pairwise logit differences for each candidate yiy_i: si=∑j=1N(Mij−Mji)s_i = \sum_{j=1}^N (M_i^j - M_j^i) Candidates are sorted in descending order of sis_i. MaxLogits is the default and best-performing aggregator.

    2. MaxWins (O(N2)\mathcal{O}(N^2) comparisons): Counts the net number of pairwise comparison victories for candidate yiy_i: si=∣{j∣Mij>0}∣+∣{j∣Mji<0}∣s_i = |\{j \mid M_i^j > 0\}| + |\{j \mid M_j^i < 0\}|

    3. Bubble Sort Selection (O(N)\mathcal{O}(N) comparisons): Identifies the top candidate efficiently with N−1N-1 comparisons:

    Input: Candidate set Y={y1,y2,…,yN}Y = \{y_1, y_2, \dots, y_N\}, Pairwise evaluator ff
    Output: Best candidate index kk
    Shuffle the candidate list YY
    Initialize k←1k \leftarrow 1
    for i=2i = 2 to NN do
        Compute Mik=f(x,yi,yk)M_i^k = f(x, y_i, y_k) and Mki=f(x,yk,yi)M_k^i = f(x, y_k, y_i)
        if Mik−Mki>0M_i^k - M_k^i > 0 then
            k←ik \leftarrow i
        end if
    end for
    return kk
  6. Knowl 6 — GENFUSER Generative Fusion Architecture

    model/method

    GENFUSER is a sequence-to-sequence model that generates an enhanced final response by fusing the top KK candidates identified by PAIRRANKER.

    • Base Model: Uses Flan-T5-XL with 3 billion parameters.

    • Input Formatting: For an input instruction xx and the top KK selected candidate responses {y(1),y(2),…,y(K)}\{y_{(1)}, y_{(2)}, \dots, y_{(K)}\} (with K=3K=3 by default), the input sequence is formed by concatenating the prompt and candidates with sentinel tokens: \text{""} x \text{ <extra_id_0> } y_{(1)} \text{ <extra_id_1> } y_{(2)} \text{ <extra_id_2> } y_{(3)} \text{""}

    • Training Objective: The model is fine-tuned to autoregressively decode the target ground truth response yy conditioned on the concatenated prompt and top candidate selections.

  7. Knowl 7 — Empirical Performance on MixInstruct Benchmark

    data/table

    Evaluation of individual LLMs, oracle selections, baseline rankers, PAIRRANKER, and the full LLM-BLENDER framework (PAIRRANKER with K=3K=3 plus GENFUSER) on the 5,000-example MixInstruct test set.

    Metrics include reference-based NLG scores (BERTScore, BARTScore, BLEURT), average rank based on ChatGPT pairwise comparisons (GPT-Rank, where lower is better across 11 LLMs), the percentage of instances matching or beating Vicuna (≥Vic\ge\text{Vic}) and OpenAssistant (≥OA\ge\text{OA}), and the proportion of instances ranking in the top 3 (Top-3).

    Methods BERTScore ↑\uparrow BARTScore ↑\uparrow BLEURT ↑\uparrow GPT-Rank ↓\downarrow ≥\ge Vic (%) ↑\uparrow ≥\ge OA (%) ↑\uparrow Top-3 (%) ↑\uparrow
    LLMs
    Open Assistant 74.68 -3.45 -0.39 3.90 62.78 N/A 51.98
    Vicuna 69.60 -3.44 -0.61 4.13 N/A 64.77 52.88
    Alpaca 71.46 -3.57 -0.53 4.62 56.70 61.35 44.46
    Baize 65.57 -3.53 -0.66 4.86 52.76 56.40 38.80
    MOSS 64.85 -3.65 -0.73 5.09 51.62 51.79 38.27
    ChatGLM 70.38 -3.52 -0.62 5.63 44.04 45.67 28.78
    Koala 63.96 -3.85 -0.84 6.76 39.93 39.01 22.55
    Dolly V2 62.26 -3.83 -0.87 6.90 33.33 31.44 16.45
    Mosaic MPT 63.21 -3.72 -0.82 7.19 30.87 30.16 16.24
    StableLM 62.47 -4.12 -0.98 8.71 21.55 19.87 7.96
    Flan-T5 64.92 -4.57 -1.23 8.81 23.89 19.93 5.32
    Analysis (Oracles)
    Oracle (BERTScore) 77.67 -3.17 -0.27 3.88 54.41 38.84 53.49
    Oracle (BLEURT) 75.02 -3.15 -0.15 3.77 55.61 45.80 55.36
    Oracle (BARTScore) 73.23 -2.87 -0.38 3.69 50.32 57.01 57.33
    Oracle (GPT-Rank) 70.32 -3.33 -0.51 1.00 100.00 100.00 100.00
    Rankers
    Random 66.36 -3.76 -0.77 6.14 37.75 36.91 29.05
    MLM-Scoring 64.77 -4.03 -0.88 7.00 33.87 30.39 21.46
    SimCLS 73.14 -3.22 -0.38 3.50 52.11 49.93 60.72
    SummaReranker 71.60 -3.25 -0.41 3.66 55.63 48.46 57.54
    PairRanker 72.97 -3.14 -0.37 3.20 54.76 57.79 65.12
    LLM-BLENDER
    PR (K=3K=3) + GF 79.09 -3.02 -0.17 3.01 70.73 77.72 68.59

    PAIRRANKER achieves an average GPT-Rank of 3.20, outperforming the best standalone LLM (OpenAssistant at 3.90). LLM-BLENDER achieves an average GPT-Rank of 3.01 and the best scores across all reference-based metrics (BERTScore 79.09, BARTScore -3.02, BLEURT -0.17), placing in the top 3 in 68.59% of instances.

  8. Knowl 8 — Correlation Analysis of Evaluators and Rankers with GPT-Rank

    data/table

    Correlation between various candidate ranking methods, automated evaluation metrics, and the oracle pairwise ranking derived from ChatGPT (GPT-Rank) on MixInstruct.

    Ranking Methods Pearson Correlation ↑\uparrow Spearman's Correlation ↑\uparrow Spearman's Footrule ↓\downarrow
    Random 0.00 0.00 48.27
    BLEU 28.70 26.92 33.57
    Rouge2 29.17 27.77 32.96
    BERTScore 32.25 30.33 33.34
    BLEURT 34.14 32.31 32.17
    BARTScore 38.49 36.76 30.93
    MLM-Scoring -0.02 -0.01 47.16
    SimCLS 39.89 38.13 29.32
    SummaReranker 41.13 39.10 29.69
    PairRanker 46.98 44.98 27.52

    Key takeaways:

    1. BARTScore shows the strongest correlation among reference-based metrics with GPT-Rank (38.49 Pearson, 36.76 Spearman), supporting its use as supervision for training PAIRRANKER.
    2. Pointwise rankers (SimCLS and SummaReranker) outperform standalone metrics but underperform pairwise modeling.
    3. PAIRRANKER achieves the highest alignment with GPT-Rank across all correlation metrics (Pearson 46.98, Spearman 44.98, Spearman's Footrule distance 27.52).
  9. Knowl 9 — PAIRRANKER Training Efficiency, Order Shuffling, and Ground Truth Augmentation

    model/method

    PAIRRANKER employs three training optimizations to handle the quadratic combination space of candidate pairs:

    1. Candidate Pair Sub-Sampling: Instead of computing all N(N−1)/2N(N-1)/2 possible candidate pairs per prompt, exactly 5 pairs are randomly sampled during training, which maintains accuracy while preserving efficiency.

    2. Ground Truth Augmentation: The target reference response yy is added directly into the candidate set YY during training so the model learns comparisons against high-quality reference answers.

    3. Candidate Order Shuffling: Because transformer positional embeddings can introduce order bias, candidate order in each pair (x,yi,yj)(x, y_i, y_j) versus (x,yj,yi)(x, y_j, y_i) is randomly shuffled during training to enforce self-consistency.

  10. Knowl 10 — PAIRRANKER Computational Overhead and Metric Limitations

    limitation

    LLM-BLENDER is subject to two main limitations:

    1. Inference Complexity: Computing the full pairwise score matrix M\mathbf{M} requires O(N2)\mathcal{O}(N^2) encoder calls for NN candidate responses. While these inferences are independent and can be executed in parallel, full evaluation increases computational latency relative to pointwise ranking. While bubble sort aggregation reduces the evaluation count to O(N)\mathcal{O}(N), it yields slightly lower ranking performance than MaxLogits.

    2. Reliance on Automated and LLM-Based Evaluation: Because large-scale human evaluation across 11 models and 110k instructions is cost-prohibitive, MixInstruct relies on ChatGPT pairwise judgments (GPT-Rank) and reference metrics (BARTScore, BERTScore, BLEURT), which may contain automated evaluation biases.

Coverage note — Omitted prior version appendix experiments on summarization (CNN/DM), machine translation (WMT18-zh-en), and constrained generation (CommonGen), as the main paper focuses exclusively on instruction-following LLM ensembling.

References

  1. 1.Anna Anioł and Marcin Pietroń. 2019. Ensemble approach for natural language question answering problem. 2019 Seventh International Symposium on Computing and Networking Workshops (CANDARW), pages 180–183.
  2. 2.Stella Rose Biderman, Hailey Schoelkopf, Quentin G. Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. ArXiv preprint, abs/2304.01373.
  3. 3.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017. Findings of the 2017 conference on machine translation (WMT17). In Proceedings of the Second Conference on Machine Translation, pages 169–214, Copenhagen, Denmark. Association for Computational Linguistics.
  4. 4.Sebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, John A. Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuan-Fang Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. ArXiv preprint, abs/2303.12712.
  5. 5.Christopher J. C. Burges. 2010. From ranknet to lambdarank to lambdamart: An overview.
  6. 6.Christopher J. C. Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Gregory N. Hullender. 2005. Learning to rank using gradient descent. In Machine Learning, Proceedings of the Twenty-Second International Conference (ICML 2005), Bonn, Germany, August 7-11, 2005, volume 119 of ACM International Conference Proceeding Series, pages 89–96. ACM.
  7. 7.Alex Cabrera and Graham Neubig. 2023. Zeno chatbot report. Blog post.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcıa, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Dıaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. ArXiv preprint, abs/2204.02311.
  10. 10.Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. ArXiv preprint, abs/2210.11416.
  11. 11.Mike Conover, Matt Hayes, Ankit Mathur, Xiangrui Meng, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  12. 12.Persi Diaconis and Ron Graham. 1977. Spearman’s footrule as a measure of disarray. Journal of the royal statistical society series b-methodological, 39:262–268.
  13. 13.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335, Dublin, Ireland. Association for Computational Linguistics.
  14. 14.Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. Koala: A dialogue model for academic research. Blog post.
  15. 15.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pretraining with gradient-disentangled embedding sharing. ArXiv preprint, abs/2111.09543.
  16. 16.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 1693–1701.
  17. 17.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  18. 18.Kevin G. Jamieson and Robert D. Nowak. 2011. Active ranking using pairwise comparisons. In Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain, pages 2240–2248.
  19. 19.LAION-AI. 2023. Open assistant. https://github.com/LAION-AI/Open-Assistant.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  21. 21.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv preprint, abs/1907.11692.
  23. 23.Yixin Liu and Pengfei Liu. 2021. SimCLS: A simple framework for contrastive learning of abstractive summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 1065–1072, Online. Association for Computational Linguistics.
  24. 24.NLP Team MosaicML. 2023. Introducing mpt-7b: A new standard for open-source, ly usable llms. Accessed: 2023-05-23.
  25. 25.Ramesh Nallapati, Bowen Zhou, Cıcero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  26. 26.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. ArXiv preprint, abs/2203.02155.
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  28. 28.Mathieu Ravaut, Shafiq Joty, and Nancy Chen. 2022a. SummaReranker: A multi-task mixture-of-experts re-ranking framework for abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4504–4524, Dublin, Ireland. Association for Computational Linguistics.
  29. 29.Mathieu Ravaut, Shafiq Joty, and Nancy Chen. 2022b. Towards summary candidates fusion. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8488–8504, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  30. 30.Omer Sagi and Lior Rokach. 2018. Ensemble learning: A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8.
  31. 31.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
  32. 32.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  33. 33.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 4603–4611. PMLR.
  34. 34.Stability-AI. 2023. Stablelm: Stability ai language models. https://github.com/stability-AI/stableLM.
  35. 35.Tianxiang Sun and Xipeng Qiu. 2023. Moss. https://github.com/OpenLMLab/MOSS.
  36. 36.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  37. 37.Jorg Tiedemann and Santhosh Thottingal. 2020a. OPUS-MT – building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pages 479–480, Lisboa, Portugal. European Association for Machine Translation.
  38. 38.Jorg Tiedemann and Santhosh Thottingal. 2020b. OPUS-MT – building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pages 479–480, Lisboa, Portugal. European Association for Machine Translation.
  39. 39.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, Aur’elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971.
  40. 40.Benyou Wang, Jiabin Niu, Liqun Ma, Yuhua Zhang, Lipeng Zhang, Jingfei Li, Peng Zhang, and Dawei Song. 2016. A chinese question answering approach integrating count-based and embedding-based features. In NLPCC/ICCPOL.
  41. 41.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. ArXiv preprint, abs/2212.10560.
  42. 42.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. ArXiv preprint, abs/2304.01196.
  43. 43.Wang Yidong, Yu Zhuohao, Zeng Zhengran, Yang Linyi, Heng Qiang, Wang Cunxiang, Chen Hao, Jiang Chaoya, Xie Rui, Wang Jindong, Xie Xing, Ye Wei, Zhang Shikun, and Zhang Yue. 2023. Pandalm: Reproducible and automated language model assessment. https://github.com/WeOpenML/PandaLM.
  44. 44.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 27263–27277.
  45. 45.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020a. PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 11328–11339. PMLR.
  46. 46.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020b. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  47. 47.Lianmin Zheng, Ying Sheng, Wei-Lin Chiang, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Chatbot arena: Benchmarking llms in the wild with elo ratings. Blog post.

Citation

MLA
Jiang, D., et al. “LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 14165–78, https://doi.org/10.18653/v1/2023.acl-long.792.
APA
Jiang, D., Ren, X., & Lin, B. Y. (2023). LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14165–14178. https://doi.org/10.18653/v1/2023.acl-long.792
Chicago
Jiang, D., X. Ren, and B. Y. Lin. 2023. “LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14165–78. https://doi.org/10.18653/v1/2023.acl-long.792.
Harvard
Jiang, D., Ren, X. and Lin, B.Y. (2023) “LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14165–14178. Available at: https://doi.org/10.18653/v1/2023.acl-long.792.
Vancouver
1. Jiang D, Ren X, Lin BY (2023) LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14165–14178

BibTeX

@inproceedings{jiang-etal-2023-llm,
    title = "{LLM}-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion",
    author = "Jiang, Dongfu  and
      Ren, Xiang  and
      Lin, Bill Yuchen",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.792/",
    doi = "10.18653/v1/2023.acl-long.792",
    pages = "14165--14178"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/