SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization
Mathieu RavautShafiq R. JotyNancy F. Chen
Proposes a multi-task mixture-of-experts re-ranking framework that jointly optimizes across multiple evaluation metrics to select superior candidate summaries from standard generation models, setting new state-of-the-art results across several standard abstractive summarization benchmarks.
Modern automated text summarization models frequently fail to output their best possible summary. Standard systems generate text word by word using search algorithms like beam search, which pick a single summary based on sequence probabilities. However, this process often discards other candidate summaries that are significantly higher in quality, coherence, and factual accuracy. Analysis indicates that the theoretical best summary among a pool of generated options—known as the oracle score—can exceed standard outputs by up to 30.5% in standard overlap metrics (ROUGE-1). This massive gap demonstrates that underlying text generation models are currently underutilized due to suboptimal candidate selection.
The article demonstrates and evaluates a second-stage re-ranking system called SummaReranker. The objective is to design a lightweight, multi-task framework that re-evaluates a pool of summary candidates generated by an underlying language model and selects the highest-quality output.
To achieve this, the authors implemented a mixture-of-experts model built on top of RoBERTa-large. The re-ranker evaluates candidate summaries jointly with the original source document and estimates the probability that a candidate is the best choice across multiple standard evaluation metrics simultaneously (including word-overlap metrics like ROUGE and embedding-based metrics like BERTScore and BARTScore). The approach was evaluated across three distinct datasets representing news and social media domains (CNN-DailyMail, XSum, and Reddit TIFU) using two state-of-the-art base language models (PEGASUS and BART).
The evaluation produced several key findings. First, SummaReranker established a new state of the art in summarization, improving standard overlap performance by 5.44% on CNN-DailyMail, 1.31% on XSum, and 9.34% on Reddit TIFU compared to standard generation. Second, combining diverse candidate generation methods (such as standard beam search and diverse beam search) consistently improved re-ranking quality. Third, base models trained on only 50% of the training data and augmented with SummaReranker outperformed models trained on 100% of the data without re-ranking. Finally, both automated recall analyses and human evaluations confirmed that the re-ranker consistently selects summaries that are more abstractive, complete, and faithful than baseline outputs.
These findings indicate that organizations deploying automated summarization can achieve substantial quality improvements without the massive computational expense of retraining foundational generation models. Training the re-ranker requires only a fraction of the compute of the primary model (approximately four days on a single graphics processing unit for standard news datasets). While scoring candidates adds a modest latency during live generation (roughly 38 milliseconds per candidate), practical trade-off analysis shows that scoring as few as six to eight candidates captures most performance gains efficiently.
For practical implementation, organizations should adopt second-stage re-ranking when summarization quality and factual consistency are critical, using a balanced pool of six to eight candidates generated via diverse search strategies. Before deploying the system on lengthy documents (such as technical manuals or legal contracts), technical teams must conduct further development, as the current model is constrained by a 512-token input limit. Users should also note that performance gains vary depending on how abstractive the baseline domain already is, showing dramatic gains on standard text but narrower improvements on highly condensed single-sentence summaries.
- Paper: PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, Jingqing Zhang et al. (2020). This paper introduces the PEGASUS abstractive summarization model that serves as the primary base sequence-to-sequence generator evaluated and re-ranked in SummaReranker.
- Paper: Discriminative Reranking for Natural Language Parsing, Michael Collins et al. (2005). This foundational work establishes the second-stage discriminative re-ranking framework on candidate outputs generated by a base model, which SummaReranker adapts for abstractive summarization.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This work analyzes the fundamental issues of exposure bias and sequence-level metric mismatch in neural generation that motivate training a second-stage candidate re-ranker.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). This paper introduces the extreme summarization (XSum) task and dataset that serves as a core benchmark for evaluating SummaReranker.
- Paper: ROUGE: A Package for Automatic Evaluation of Summaries, Chin-Yew Lin (2004). This paper presents the ROUGE metric suite used as the primary evaluation standard and multi-task optimization target across candidate summaries.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). This study demonstrates how pre-trained Transformer encoders can be adapted for abstractive summarization, providing foundational architecture for modern sequence-to-sequence base models.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). This paper examines factual inconsistency and hallucination issues in abstractive summaries that candidate selection and re-ranking aim to address.
- Paper: Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents, Weiwei Sun et al. (2023). This paper explores how large language models can act as zero-shot listwise and pairwise re-ranking agents, extending candidate re-ranking beyond specialized discriminative architectures.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). This work extends post-generation quality improvement by replacing discrete candidate re-ranking with iterative self-feedback and refinement within language models.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). This work unifies context candidate ranking with generation in a single instruction-tuned model, continuing the evolution of ranking in generative natural language processing pipelines.
