BRIO: Bringing Order to Abstractive Summarization
Yixin LiuPengfei LiuDragomir R. RadevGraham Neubig
Introduces a contrastive training paradigm that coordinates model probabilities with candidate summary quality to overcome exposure bias and achieve state-of-the-art abstractive summarization performance on CNN/DailyMail and XSum.
Automated text summarization systems are increasingly critical for organizations managing large volumes of textual data, yet standard training techniques often produce models that degrade during deployment. Conventional systems rely on maximum likelihood estimation, a training method that assumes only the single human reference summary is acceptable. This creates exposure bias, leaving models poorly equipped to assess their own candidate outputs when slight deviations occur during generation. Preliminary findings demonstrate that standard baselines assign a higher probability to better-quality summaries in only about 55% of comparisons, indicating that model confidence is poorly aligned with actual output quality.
The article introduces and evaluates BRIO, a training paradigm designed to coordinate abstractive summarization models so that the probabilities they assign to candidate summaries directly reflect summary quality. By combining standard token-level generation training with sequence-level contrastive learning, the approach assigns the neural model a dual role: generating candidate summaries autoregressively and ranking candidate outputs based on quality metrics without requiring human references at inference.
The authors evaluated the framework across three standard benchmark news datasets—CNN/DailyMail, XSum, and New York Times—using established foundation models, namely BART and PEGASUS. Model performance was measured using standard evaluation metrics, including word-overlap scores (ROUGE) and semantic similarity measures (BERTScore), alongside assessments of calibration error, rank correlation, and few-shot efficiency.
The experimental results deliver several key findings. First, the proposed multi-task model achieved new state-of-the-art results, reaching a 47.78 ROUGE-1 score on CNN/DailyMail and 49.07 on XSum, significantly outperforming standard baselines and complex multi-model pipelines. Second, the framework dramatically improved ranking accuracy, boosting the frequency with which the model preferred higher-quality candidates from 54.80% to 79.63%. Third, unlike conventional systems whose output quality deteriorates when considering more candidates during search, the proposed model scaled effectively with broader search widths. Fourth, the model substantially improved token calibration, reducing expected calibration error by roughly one-third to one-half across datasets and effectively suppressing dataset noise, such as irrelevant web artifacts like 'click here' links. Finally, few-shot experiments showed that fine-tuning on as few as 100 to 1,000 samples yielded measurable quality improvements.
These findings indicate that aligning generation probabilities with output quality yields more dependable, higher-fidelity automated summaries while retaining standard model architectures. Because the framework achieves superior performance without requiring separate reranking networks or external extractors, it simplifies deployment architectures and reduces pipeline complexity. Furthermore, its rapid convergence—often requiring less than a single training epoch—lowers the computational overhead and financial cost of domain-specific model fine-tuning.
Organizations deploying automated summarization should consider adopting contrastive multi-task fine-tuning for existing sequence generation pipelines. The article demonstrates that the approach is effective whether using lexical overlap metrics like ROUGE or semantic metrics like BERTScore as the optimization target. For resource-constrained settings, teams can implement few-shot fine-tuning on modest sample sizes to capture meaningful performance gains without retraining large models from scratch.
Certain limitations remain. The candidate generation step requires initial computational overhead to produce diverse hypotheses across large training sets. Additionally, while token calibration improved, language models remain generally overconfident in their predictions, and abstractive performance decreases when source articles require highly novel phrasing. Overall, confidence in the reported improvements is high across standard news benchmarks, but stakeholders should conduct domain-specific pilot testing on specialized technical or non-news text before broad production rollout.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). This work establishes the foundational paradigm and challenges of mitigating exposure bias and metric mismatch in abstractive summarization using sequence-level reinforcement learning objectives.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). It introduces sequence-level training to address the discrepancy between maximum likelihood estimation and full-sequence generation metrics in neural sequence models.
- Paper: PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, Jingqing Zhang et al. (2020). It provides the underlying pre-trained sequence-to-sequence model and benchmark baselines commonly utilized and enhanced by contrastive ranking frameworks like BRIO.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). It demonstrates how to optimize summarization models directly against learned quality rankings and preference models rather than deterministic reference targets.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). It introduces standard architectural benchmarks and core evaluation settings on the CNN/Daily Mail dataset for abstractive summarization.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). It introduces the XSum dataset and defines the task formulation for extreme abstractive summarization evaluated extensively in BRIO.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). It establishes standard methodologies for adapting pre-trained contextual encoders to both extractive and abstractive document summarization.
- Paper: ROUGE: A Package for Automatic Evaluation of Summaries, Chin-Yew Lin (2004). It defines the ROUGE evaluation metrics that serve as the candidate ranking criteria and primary quality benchmarks in the source paper.
- Paper: SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization, Mathieu Ravaut et al. (2022). This paper extends candidate-based quality scoring in abstractive summarization by designing a dedicated multi-task mixture-of-experts re-ranking framework across multiple evaluation metrics.
- Paper: Learning to Rank in Generative Retrieval, Yongqi Li et al. (2024). It adapts sequence-level candidate ranking losses to generative retrieval to resolve the objective mismatch between text generation and passage ranking.
- Paper: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Rishabh Agarwal et al. (2024). It builds upon candidate generation dynamics by distilling language models on their own generated output distributions to mitigate exposure bias.
- Paper: Theoretical guarantees on the best-of-n alignment policy, Ahmad Beirami et al. (2025). It provides theoretical guarantees and distribution drift analyses for sampling and selecting top candidates among candidate generations under reward scoring.
- Paper: Human Alignment of Large Language Models through Online Preference Optimisation, Daniele Calandriello et al. (2024). It extends candidate preference optimization to online contrastive updates in text generation and summarization.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). It applies unified ranking and generation objectives directly within instruction-tuned language models for retrieval-augmented generation.
