BRIO: Bringing Order to Abstractive Summarization

Yixin LiuPengfei LiuDragomir R. RadevGraham Neubig

article2022ACL337 citations

Introduces a contrastive training paradigm that coordinates model probabilities with candidate summary quality to overcome exposure bias and achieve state-of-the-art abstractive summarization performance on CNN/DailyMail and XSum.

Listen

Automated text summarization systems are increasingly critical for organizations managing large volumes of textual data, yet standard training techniques often produce models that degrade during deployment. Conventional systems rely on maximum likelihood estimation, a training method that assumes only the single human reference summary is acceptable. This creates exposure bias, leaving models poorly equipped to assess their own candidate outputs when slight deviations occur during generation. Preliminary findings demonstrate that standard baselines assign a higher probability to better-quality summaries in only about 55% of comparisons, indicating that model confidence is poorly aligned with actual output quality.

The article introduces and evaluates BRIO, a training paradigm designed to coordinate abstractive summarization models so that the probabilities they assign to candidate summaries directly reflect summary quality. By combining standard token-level generation training with sequence-level contrastive learning, the approach assigns the neural model a dual role: generating candidate summaries autoregressively and ranking candidate outputs based on quality metrics without requiring human references at inference.

The authors evaluated the framework across three standard benchmark news datasets—CNN/DailyMail, XSum, and New York Times—using established foundation models, namely BART and PEGASUS. Model performance was measured using standard evaluation metrics, including word-overlap scores (ROUGE) and semantic similarity measures (BERTScore), alongside assessments of calibration error, rank correlation, and few-shot efficiency.

The experimental results deliver several key findings. First, the proposed multi-task model achieved new state-of-the-art results, reaching a 47.78 ROUGE-1 score on CNN/DailyMail and 49.07 on XSum, significantly outperforming standard baselines and complex multi-model pipelines. Second, the framework dramatically improved ranking accuracy, boosting the frequency with which the model preferred higher-quality candidates from 54.80% to 79.63%. Third, unlike conventional systems whose output quality deteriorates when considering more candidates during search, the proposed model scaled effectively with broader search widths. Fourth, the model substantially improved token calibration, reducing expected calibration error by roughly one-third to one-half across datasets and effectively suppressing dataset noise, such as irrelevant web artifacts like 'click here' links. Finally, few-shot experiments showed that fine-tuning on as few as 100 to 1,000 samples yielded measurable quality improvements.

These findings indicate that aligning generation probabilities with output quality yields more dependable, higher-fidelity automated summaries while retaining standard model architectures. Because the framework achieves superior performance without requiring separate reranking networks or external extractors, it simplifies deployment architectures and reduces pipeline complexity. Furthermore, its rapid convergence—often requiring less than a single training epoch—lowers the computational overhead and financial cost of domain-specific model fine-tuning.

Organizations deploying automated summarization should consider adopting contrastive multi-task fine-tuning for existing sequence generation pipelines. The article demonstrates that the approach is effective whether using lexical overlap metrics like ROUGE or semantic metrics like BERTScore as the optimization target. For resource-constrained settings, teams can implement few-shot fine-tuning on modest sample sizes to capture meaningful performance gains without retraining large models from scratch.

Certain limitations remain. The candidate generation step requires initial computational overhead to produce diverse hypotheses across large training sets. Additionally, while token calibration improved, language models remain generally overconfident in their predictions, and abstractive performance decreases when source articles require highly novel phrasing. Overall, confidence in the reported improvements is high across standard news benchmarks, but stakeholders should conduct domain-specific pilot testing on specialized technical or non-news text before broad production rollout.

Cover for BRIO: Bringing Order to Abstractive Summarization

Abstract

Abstractive summarization models are commonly trained using maximum likelihood estimation, which assumes a deterministic (one-point) target distribution in which an ideal model will assign all the probability mass to the reference summary. This assumption may lead to performance degradation during inference, where the model needs to compare several system-generated (candidate) summaries that have deviated from the reference summary. To address this problem, we propose a novel training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality. Our method achieves a new state-of-the-art result on the CNN/DailyMail (47.78 ROUGE-1) and XSum (49.07 ROUGE-1) datasets. Further analysis also shows that our model can estimate probabilities of candidate summaries that are more correlated with their level of quality.¹

Table of Contents

  • 1 Introduction
  • 2 Neural Abstractive Summarization
  • 3 Coordinating Abstractive Models
  • 4 Related Work
  • 5 Experiments
  • 5.1 Experimental Settings
  • 5.2 Results
  • 5.3 Analysis
  • 5.4 Token-level Calibration
  • 5.5 Few-shot Fine-tuning
  • 5.6 Case Study on CNNDM
  • 6 Conclusion and Future Work
  • Acknowledgements
  • References
  • A Datasets Statistics
  • B Implementation Details
  • C Details of Few-shot Fine-tuning

Knowls

  1. Knowl 1 — Non-Deterministic Target Distribution for Sequence-Level Coordination

    assumption

    Standard maximum likelihood estimation (MLE) training assumes a deterministic target distribution where all probability mass is concentrated on the reference summary S∗S^*. To enable an abstractive summarization model gg to distinguish the relative quality among generated candidate outputs when deviations occur during autoregressive decoding, the target distribution ptrue†(S∣D)p_{\text{true}}^\dagger(S \mid D) conditioned on document DD is redefined as a non-deterministic distribution over candidate summaries S\mathcal{S}:

    {ptrue†(S∣D)=1−βfor S=S∗∑S∈S∖{S∗}ptrue†(S∣D)=βfor S≠S∗ptrue†(Si∣D)>ptrue†(Sj∣D)∀Si,Sj∈S^, M(Si)>M(Sj)\begin{cases} p_{\text{true}}^\dagger(S \mid D) = 1 - \beta & \text{for } S = S^* \\ \sum_{S \in \mathcal{S} \setminus \{S^*\}} p_{\text{true}}^\dagger(S \mid D) = \beta & \text{for } S \neq S^* \\ p_{\text{true}}^\dagger(S_i \mid D) > p_{\text{true}}^\dagger(S_j \mid D) & \forall S_i, S_j \in \hat{\mathcal{S}}, \, M(S_i) > M(S_j) \end{cases}

    Here, β∈(0,1)\beta \in (0, 1) denotes the marginal probability assigned to non-reference candidates, S^\hat{\mathcal{S}} represents the set of top candidate summaries generated via diverse beam search, and M(S)M(S) is an automatic sequence quality evaluation metric (such as ROUGE or BERTScore) computed relative to the ground-truth reference S∗S^*.

  2. Knowl 2 — Sequence-Level Margin-Based Contrastive Ranking Loss

    equation

    To coordinate the model-assigned sequence probabilities with evaluation metric quality, a contrastive ranking loss is defined over pairs of candidate summaries generated by an abstractive model. For a candidate summary S=(s1,…,sl)S = (s_1, \dots, s_l), its length-normalized estimated log-probability f(S)f(S) under model parameters θ\theta given source document DD is:

    f(S)=∑t=1llog⁡pgθ(st∣D,S<t;θ)∣S∣αf(S) = \frac{\sum_{t=1}^{l} \log p_{g_\theta}(s_t \mid D, S_{<t}; \theta)}{|S|^\alpha}

    where α>0\alpha > 0 is the length penalty hyperparameter. Let {S1,S2,…,SK}\{S_1, S_2, \dots, S_K\} be a set of candidate summaries sorted in descending order of an evaluation metric MM, such that M(Si,S∗)>M(Sj,S∗)M(S_i, S^*) > M(S_j, S^*) for all i<ji < j. The contrastive ranking loss Lctr\mathcal{L}_{\text{ctr}} is:

    Lctr=∑i=1K∑j>imax⁡(0,f(Sj)−f(Si)+λij)\mathcal{L}_{\text{ctr}} = \sum_{i=1}^{K} \sum_{j > i} \max\left(0, f(S_j) - f(S_i) + \lambda_{ij}\right)

    where the rank-proportional margin λij\lambda_{ij} is defined as:

    λij=(j−i)⋅λ\lambda_{ij} = (j - i) \cdot \lambda

    with base margin hyperparameter λ>0\lambda > 0.

  3. Knowl 3 — BRIO Multi-Task Fine-Tuning Objective

    model/method

    The BRIO (BRInging Order) training framework trains a sequence-to-sequence model to jointly serve as an autoregressive generator and a reference-free sequence evaluator. Fine-tuning with the sequence-level contrastive loss alone damages the autoregressive generation capability because token-level distribution structure is lost. To preserve token-level prediction accuracy while enforcing sequence-level rank coordination, the multi-task loss Lmul\mathcal{L}_{\text{mul}} combines token-level cross-entropy loss Lxent\mathcal{L}_{\text{xent}} on the ground-truth summary S∗S^* of length ll with the contrastive loss Lctr\mathcal{L}_{\text{ctr}} over generated candidate summaries:

    Lmul=Lxent+γLctr\mathcal{L}_{\text{mul}} = \mathcal{L}_{\text{xent}} + \gamma \mathcal{L}_{\text{ctr}}

    where γ>0\gamma > 0 controls the weight of the contrastive objective, and:

    Lxent=−∑j=1l∑sptrue(s∣D,S<j∗)log⁡pgθ(s∣D,S<j∗;θ)\mathcal{L}_{\text{xent}} = - \sum_{j=1}^{l} \sum_{s} p_{\text{true}}(s \mid D, S^*_{<j}) \log p_{g_\theta}(s \mid D, S^*_{<j}; \theta)

    In this framework, BRIO-Ctr denotes the model variant trained only with Lctr\mathcal{L}_{\text{ctr}} (used strictly as a second-stage reranker over pre-generated candidates), whereas BRIO-Mul is trained with Lmul\mathcal{L}_{\text{mul}} and generates summaries end-to-end via standard autoregressive decoding.

  4. Knowl 4 — Summarization Benchmark Performance on CNN/DailyMail, XSum, and NYT

    data/table

    BRIO-Mul and BRIO-Ctr achieve state-of-the-art results across three benchmark datasets (CNN/DailyMail, XSum, and NYT) evaluated with ROUGE-1 (R-1), ROUGE-2 (R-2), and ROUGE-L (R-L) F1F_1 scores:

    System R-1 R-2 R-L
    CNN/DailyMail
    BART (Lewis et al., 2020) 44.29 21.17 41.09
    PEGASUS (Zhang et al., 2020) 44.17 21.47 41.11
    GSum (Dou et al., 2021) 45.94 22.32 42.48
    SimCLS (Liu and Liu, 2021) 46.67 22.15 43.54
    BRIO-Ctr 47.28 22.93 44.15
    BRIO-Mul 47.78 23.55 44.57
    XSum
    BART (Lewis et al., 2020) 45.14 22.27 37.25
    PEGASUS (Zhang et al., 2020) 47.46 24.69 39.53
    GSum (Dou et al., 2021) 45.40 21.89 36.67
    SimCLS (Liu and Liu, 2021) 47.61 24.57 39.44
    BRIO-Ctr 48.13 25.13 39.84
    BRIO-Mul 49.07 25.59 40.40
    NYT
    BART 55.78 36.61 52.60
    BRIO-Ctr 55.98 36.54 52.51
    BRIO-Mul 57.75 38.64 54.54

    BRIO-Ctr outperforms SimCLS on CNN/DailyMail (47.28 vs. 46.67 R-1) and XSum (48.13 vs. 47.61 R-1) by reusing the same Seq2Seq architecture for both candidate generation and reranking rather than training an auxiliary RoBERTa model. BRIO-Mul outperforms all prior methods without requiring external guidance representations as in GSum.

  5. Knowl 5 — Beam Search Width Scaling Under Sequence-Level Coordination

    empirical result

    Standard MLE-trained sequence-to-sequence models suffer performance degradation when inference beam width is increased, because expanding the search space introduces candidates of lower quality that the uncoordinated generator fails to rank below better candidates. In contrast, BRIO-Mul coordinates model probability with candidate quality, allowing performance to improve monotonically as beam width expands on CNN/DailyMail:

    Beams BART BRIO-Mul
    R-1 R-2 R-1 R-2
    4 44.29 21.17 47.78 23.55
    10 43.83 20.76 47.98 23.81
    20 43.53 20.49 48.07 23.92
    50 43.06 20.05 48.18 24.01
    100 42.79 19.76 48.23 24.09

    While BART degrades from 44.29 R-1 at beam width 4 down to 42.79 R-1 at beam width 100, BRIO-Mul gains performance from 47.78 up to 48.23 R-1.

  6. Knowl 6 — Iterative Candidate Generation and Fine-Tuning Loop (BRIO-Loop)

    model/method

    Because BRIO-Mul retains autoregressive generative ability, candidate summaries can be iteratively sampled from a trained BRIO-Mul model and used to construct a refreshed training set for an additional round of contrastive fine-tuning. This self-training variant, termed BRIO-Loop, outperforms the initial BRIO-Mul model on CNN/DailyMail:

    System R-1 R-2 R-L
    BART baseline 44.29 21.17 41.09
    BRIO-Mul 47.78 23.55 44.57
    BRIO-Loop 48.01 23.80 44.67

    BRIO-Loop achieves statistically significant improvements (p<0.01p < 0.01) over BART and reaches its peak validation performance rapidly during the second fine-tuning stage.

  7. Knowl 7 — Model-Probability and Quality-Score Rank Correlation

    empirical result

    The rank correlation between a model's length-normalized estimated log-probability f(S)f(S) and summary quality (ROUGE-1 against the reference) is evaluated using Spearman's rank correlation over 16 candidate summaries per sample on CNN/DailyMail. Evaluated on candidates generated by the models themselves ('Own') and on candidates generated independently by a pre-trained PEGASUS model ('PEGASUS'):

    Model Own PEGASUS
    BART baseline 0.0470 0.1205
    BRIO-Mul 0.1839 0.2768

    BRIO-Mul demonstrates substantially higher rank correlation than BART both on self-generated candidates (0.1839 vs. 0.0470) and on out-of-distribution candidates produced by PEGASUS (0.2768 vs. 0.1205), demonstrating improved general calibration of sequence-level scores to summary quality.

  8. Knowl 8 — Token-Level Calibration Improvement via Sequence-Level Contrastive Training

    empirical result

    Enforcing sequence-level coordination via multi-task contrastive learning directly reduces token-level calibration error during inference. Token calibration is measured by Expected Calibration Error (ECE) across MM equal-width confidence buckets BmB_m:

    ECE=∑m=1M∣Bm∣n∣acc(Bm)−conf(Bm)∣\text{ECE} = \sum_{m=1}^{M} \frac{|B_m|}{n} |\text{acc}(B_m) - \text{conf}(B_m)|

    where token correctness labels on system-generated summaries are assigned using the Tercom toolkit against reference summaries. On CNN/DailyMail and XSum test sets:

    Dataset System ECE ↓\downarrow Accuracy ↑\uparrow Confidence
    CNN/DailyMail BART baseline 0.4097 0.3711 0.7365
    BRIO-Mul 0.2719 0.4271 0.6652
    XSum PEGASUS baseline 0.2369 0.4688 0.6990
    BRIO-Mul 0.1423 0.4744 0.5881

    While standard MLE models display severe over-confidence (average confidence exceeds actual token accuracy by 30-36%), BRIO-Mul lowers ECE by 33.6% on CNN/DailyMail and 39.9% on XSum while simultaneously improving prediction accuracy.

  9. Knowl 9 — Generalization Across Metrics and Abstractive Novel N-Grams

    empirical result

    BRIO is metric-agnostic: candidate summaries can be ordered using semantic metrics such as BERTScore (BS) instead of ROUGE in the contrastive loss Lctr\mathcal{L}_{\text{ctr}}. On CNN/DailyMail, training BRIO-Mul with candidates ordered by BERTScore (BRIO-Mul (B)) improves both ROUGE and BERTScore over baseline BART, while training with ROUGE (BRIO-Mul (R)) also boosts BERTScore:

    System R-1 R-2 R-L BERTScore
    BART 44.29 21.17 41.09 27.38
    BRIO-Mul (R) 47.78 23.55 44.57 32.11
    BRIO-Mul (B) 47.53 23.22 44.37 32.59

    Furthermore, BRIO-Mul exhibits higher abstractiveness by generating higher ratios of novel unigrams (0.0262 vs. 0.0101 for BART) and novel bigrams (0.2381 vs. 0.0924 for BART). The performance advantage of BRIO-Mul over BART increases monotonically on document subsets with higher reference bigram novelty:

    Novelty(D,S∗)=∑g∈GS∗1(g∉GD)∣GS∗∣\text{Novelty}(D, S^*) = \frac{\sum_{g \in G_{S^*}} \mathbf{1}(g \notin G_D)}{|G_{S^*}|}

    where GS∗G_{S^*} and GDG_D are the sets of bigrams in the reference summary and source document, respectively.

  10. Knowl 10 — Few-Shot Fine-Tuning with Sequence Contrastive Learning

    empirical result

    BRIO contrastive fine-tuning yields substantial performance improvements in low-resource and few-shot scenarios without requiring candidate generation across the complete training dataset. Training BRIO-Few on randomly sampled subsets (100 training examples for CNN/DailyMail, 1000 training examples for XSum) using Adam with learning rate 1×10−61 \times 10^{-6}:

    Dataset System R-1 R-2 R-L
    CNN/DailyMail BART baseline 44.29 21.17 41.09
    BRIO-Few (100) 45.81 21.91 42.61
    XSum PEGASUS baseline 47.46 24.69 39.53
    BRIO-Few (1000) 47.95 24.89 39.71

    Fine-tuning on just 100 CNN/DailyMail examples boosts R-1 by +1.52 points over the fully trained MLE BART baseline, confirming that sequence-level candidate ranking provides high sample efficiency.

Coverage note — None was omitted; all key theoretical formulations, loss functions, benchmark evaluations, ablation and loop analyses, beam search behaviors, calibration, and few-shot results are captured.

References

  1. 1.Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio. 2016. An actor-critic algorithm for sequence prediction. CoRR, abs/1607.07086.
  2. 2.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, page 1171–1179, Cambridge, MA, USA. MIT Press.
  3. 3.Shuyang Cao and Lu Wang. 2021. CLIFF: Contrastive learning for improving faithfulness and factuality in abstractive summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6633–6649, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1597–1607. PMLR.
  5. 5.Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. On the properties of neural machine translation: Encoder–decoder approaches. In Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, pages 103–111, Doha, Qatar. Association for Computational Linguistics.
  6. 6.Woon Sang Cho, Yizhe Zhang, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao, Mengdi Wang, and Bill Dolan. 2021. Contrastive multi-document question generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 12–30, Online. Association for Computational Linguistics.
  7. 7.Sumit Chopra, Michael Auli, and Alexander M. Rush. 2016. Abstractive sentence summarization with attentive recurrent neural networks. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 93–98, San Diego, California. Association for Computational Linguistics.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021. GSum: A general framework for guided neural abstractive summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4830–4842, Online. Association for Computational Linguistics.
  10. 10.Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018. Classical structured prediction losses for sequence to sequence learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 355–364, New Orleans, Louisiana. Association for Computational Linguistics.
  11. 11.Alexander Fabbri, Simeng Han, Haoyuan Li, Haoran Li, Marjan Ghazvininejad, Shafiq Joty, Dragomir Radev, and Yashar Mehdad. 2021. Improving zero and few-shot abstractive summarization with intermediate fine-tuning and data augmentation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 704–717, Online. Association for Computational Linguistics.
  12. 12.Kevin Gimpel and Noah A. Smith. 2010. Softmax-margin CRFs: Training log-linear models with cost functions. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 733–736, Los Angeles, California. Association for Computational Linguistics.
  13. 13.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR.
  14. 14.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 2, CVPR ’06, page 1735–1742, USA. IEEE Computer Society.
  15. 15.Ralf Herbrich, Thore Graepel, and Klaus Obermayer. 1999. Support vector learning for ordinal regression. In In International Conference on Artificial Neural Networks, pages 97–102.
  16. 16.Karl Moritz Hermann, Tomáš Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, page 1693–1701, Cambridge, MA, USA. MIT Press.
  17. 17.Mark Hopkins and Jonathan May. 2011. Tuning as ranking. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, pages 1352–1362, Edinburgh, Scotland, UK. Association for Computational Linguistics.
  18. 18.Chris Kedzie, Kathleen McKeown, and Hal Daumé III. 2018. Content selection in deep learning models of summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1818–1828, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.Huda Khayrallah, Brian Thompson, Matt Post, and Philipp Koehn. 2020. Simulated multiple reference training improves low-resource machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 82–89, Online. Association for Computational Linguistics.
  20. 20.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  21. 21.Aviral Kumar and Sunita Sarawagi. 2019. Calibration of encoder decoder models for neural machine translation. CoRR, abs/1903.00802.
  22. 22.Ann Lee, Michael Auli, and Marc’Aurelio Ranzato. 2021a. Discriminative reranking for neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7250–7264, Online. Association for Computational Linguistics.
  23. 23.Seanie Lee, Dong Bok Lee, and Sung Ju Hwang. 2021b. Contrastive learning with adversarial perturbations for conditional text generation. In International Conference on Learning Representations.
  24. 24.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  25. 25.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. Deep reinforcement learning for dialogue generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Austin, Texas. Association for Computational Linguistics.
  26. 26.Siyao Li, Deren Lei, Pengda Qin, and William Yang Wang. 2019. Deep reinforcement learning with distributional semantic rewards for abstractive summarization. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6038–6044, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  28. 28.Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaichen Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, and Graham Neubig. 2021a. ExplainaBoard: An explainable leaderboard for NLP. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 280–289, Online. Association for Computational Linguistics.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  30. 30.Yixin Liu, Zi-Yi Dou, and Pengfei Liu. 2021b. RefSum: Refactoring neural summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1437–1448, Online. Association for Computational Linguistics.
  31. 31.Yixin Liu and Pengfei Liu. 2021. SimCLS: A simple framework for contrastive learning of abstractive summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 1065–1072, Online. Association for Computational Linguistics.
  32. 32.Tomoya Mizumoto and Yuji Matsumoto. 2016. Discriminative reranking for grammatical error correction with statistical machine translation. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1133–1138, San Diego, California. Association for Computational Linguistics.
  33. 33.Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. 2019. When does label smoothing help? In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  34. 34.Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, AAAI’15, page 2901–2907. AAAI Press.
  35. 35.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çaglar Gulçehre, and Bing Xiang. 2016. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  36. 36.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Mohammad Norouzi, Samy Bengio, zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans. 2016. Reward augmented maximum likelihood for neural structured prediction. In Advances in Neural Information Processing Systems, volume 29, pages 1723–1731. Curran Associates, Inc.
  38. 38.Franz Josef Och. 2003. Minimum error rate training in statistical machine translation. In Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, pages 160–167, Sapporo, Japan. Association for Computational Linguistics.
  39. 39.Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alex Fraser, Shankar Kumar, Libin Shen, David Smith, Katherine Eng, Viren Jain, Zhen Jin, and Dragomir Radev. 2004. A smorgasbord of features for statistical machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 161–168, Boston, Massachusetts, USA. Association for Computational Linguistics.
  40. 40.Xiao Pan, Mingxuan Wang, Liwei Wu, and Lei Li. 2021. Contrastive learning for many-to-many multilingual neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 244–258, Online. Association for Computational Linguistics.
  41. 41.Richard Yuanzhe Pang and He He. 2021. Text generation by learning from demonstrations. In International Conference on Learning Representations.
  42. 42.Romain Paulus, Caiming Xiong, and Richard Socher. 2018. A deep reinforced model for abstractive summarization. In International Conference on Learning Representations.
  43. 43.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  44. 44.Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016. Sequence level training with recurrent neural networks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  45. 45.Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 379–389, Lisbon, Portugal. Association for Computational Linguistics.
  46. 46.Evan Sandhaus. 2008. The New York Times Annotated Corpus. LDC corpora. Linguistic Data Consortium.
  47. 47.Timo Schick and Hinrich Schütze. 2021. Few-shot text generation with natural language instructions. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 390–402, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  48. 48.Libin Shen, Anoop Sarkar, and Franz Josef Och. 2004. Discriminative reranking for machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 177–184, Boston, Massachusetts, USA. Association for Computational Linguistics.
  49. 49.Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016. Minimum risk training for neural machine translation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1683–1692, Berlin, Germany. Association for Computational Linguistics.
  50. 50.Felix Stahlberg and Bill Byrne. 2019. On NMT search errors and model errors: Cat got your tongue? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3356–3362, Hong Kong, China. Association for Computational Linguistics.
  51. 51.Shichao Sun and Wenjie Li. 2021. Alleviating exposure bias via contrastive learning for abstractive text summarization. CoRR, abs/2108.11846.
  52. 52.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2, NIPS’14, page 3104–3112, Cambridge, MA, USA. MIT Press.
  53. 53.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. 2016. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, Los Alamitos, CA, USA. IEEE Computer Society.
  54. 54.Ben Taskar, Carlos Guestrin, and Daphne Koller. 2004. Max-margin markov networks. In Advances in Neural Information Processing Systems, volume 16. MIT Press.
  55. 55.Yui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, and Yusuke Miyao. 2020. Learning with contrastive examples for data-to-text generation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2352–2362, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  56. 56.Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse beam search for improved description of complex scenes. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1).
  57. 57.Xiaojun Wan, Ziqiang Cao, Furu Wei, Sujian Li, and M. Zhou. 2015. Multi-document summarization via discriminative summary reranking. ArXiv, abs/1507.02062.
  58. 58.Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020. On the inference calibration of neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3070–3079, Online. Association for Computational Linguistics.
  59. 59.John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019. Beyond BLEU:training neural machine translation with semantic similarity. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4344–4355, Florence, Italy. Association for Computational Linguistics.
  60. 60.Sam Wiseman and Alexander M. Rush. 2016. Sequence-to-sequence learning as beam-search optimization. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1296–1306, Austin, Texas. Association for Computational Linguistics.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  62. 62.Shusheng Xu, Xingxing Zhang, Yi Wu, and Furu Wei. 2021. Sequence level contrastive learning for text summarization. CoRR, abs/2109.03481.
  63. 63.Zonghan Yang, Yong Cheng, Yang Liu, and Maosong Sun. 2019. Reducing word omission errors in neural machine translation: A contrastive learning approach. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6191–6196, Florence, Italy. Association for Computational Linguistics.
  64. 64.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation. In Thirty-Fifth Conference on Neural Information Processing Systems.
  65. 65.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  66. 66.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  67. 67.Wen Zhang, Yang Feng, Fandong Meng, Di You, and Qun Liu. 2019. Bridging the gap between training and inference for neural machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4334–4343, Florence, Italy. Association for Computational Linguistics.
  68. 68.Ming Zhong, Pengfei Liu, Yiran Chen, Danqing Wang, Xipeng Qiu, and Xuanjing Huang. 2020. Extractive summarization as text matching. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6197–6208, Online. Association for Computational Linguistics.

Citation

MLA
Liu, Y., et al. “BRIO: Bringing Order to Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2890–903, https://doi.org/10.18653/v1/2022.acl-long.207.
APA
Liu, Y., Liu, P., Radev, D., & Neubig, G. (2022). BRIO: Bringing Order to Abstractive Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2890–2903. https://doi.org/10.18653/v1/2022.acl-long.207
Chicago
Liu, Y., P. Liu, D. Radev, and G. Neubig. 2022. “BRIO: Bringing Order to Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2890–2903. https://doi.org/10.18653/v1/2022.acl-long.207.
Harvard
Liu, Y. et al. (2022) “BRIO: Bringing Order to Abstractive Summarization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2890–2903. Available at: https://doi.org/10.18653/v1/2022.acl-long.207.
Vancouver
1. Liu Y, Liu P, Radev D, Neubig G (2022) BRIO: Bringing Order to Abstractive Summarization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2890–2903

BibTeX

@inproceedings{liu-etal-2022-brio,
    title = "{BRIO}: Bringing Order to Abstractive Summarization",
    author = "Liu, Yixin  and
      Liu, Pengfei  and
      Radev, Dragomir  and
      Neubig, Graham",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.207/",
    doi = "10.18653/v1/2022.acl-long.207",
    pages = "2890--2903"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/