PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
Jingqing ZhangYao ZhaoMohammad SalehPeter J. Liu
Proposes a gap-sentence pre-training objective specifically designed for abstractive text summarization, setting state-of-the-art performance across 12 diverse datasets while demonstrating remarkable sample efficiency with as few as 1,000 fine-tuning examples.
The article addresses the challenge of improving abstractive text summarization through pre-training of Transformer models. While general self-supervised pre-training has advanced many NLP tasks, objectives specifically suited to summarization remained unexplored, and evaluations lacked breadth across domains.
This work sets out to develop and test a new pre-training objective called Gap Sentences Generation for encoder-decoder Transformers.
The approach involves pre-training on large corpora by masking important sentences and reconstructing them from the rest of the document. Models were trained on C4 and HugeNews datasets, with ablations performed on smaller models before scaling to 568 million parameters. Performance was assessed on 12 diverse summarization datasets using ROUGE metrics and human evaluations.
Key findings show that the best PEGASUS model achieves state-of-the-art results on all 12 tasks. It also delivers strong performance in low-resource scenarios, exceeding prior benchmarks on six datasets using just 1,000 examples. Human judges rated its summaries as comparable to reference ones on several datasets.
These results indicate that task-aligned pre-training objectives can substantially boost summarization quality and adaptability, particularly when labeled data is limited. Alignment between pre-training and downstream domains further enhances transfer.
Next steps include exploring larger models, handling longer inputs, and applying similar ideas to other generation tasks.
Limitations involve potential data overlap effects, though minimal, and the focus on specific model architectures. Confidence is high given extensive experiments and human validation.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Introduces denoising sequence-to-sequence pre-training using bidirectional encoders and autoregressive decoders, establishing the foundational architecture that PEGASUS customizes for summarization.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Introduces the unified text-to-text encoder-decoder framework and the large-scale C4 corpus, both of which serve as core foundations for PEGASUS's pre-training pipeline.
- Paper: SpanBERT: Improving Pre-training by Representing and Predicting Spans, Mandar Joshi et al. (2019). Demonstrates the efficacy of masking contiguous multi-token spans during pre-training, directly motivating PEGASUS's sentence-level gap-masking objective.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). Establishes standard neural abstractive summarization baselines, datasets, and evaluation metrics that PEGASUS benchmarks against.
- Paper: Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond, Ramesh Nallapati et al. (2016). Provides fundamental techniques and benchmark datasets (like CNN/Daily Mail) for sequence-to-sequence abstractive summarization evaluated extensively in PEGASUS.
- Paper: A Neural Attention Model for Abstractive Sentence Summarization, Alexander M. Rush et al. (2015). Pioneered neural attention-based encoder-decoder modeling for abstractive summarization, setting the task framing that modern pre-trained models advance.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). Established the masked-language-modeling paradigm for self-supervised Transformer pre-training that PEGASUS adapts to full sentence generation.
- Paper: RoBERTa: A Robustly Optimized BERT Pretraining Approach, Yinhan Liu et al. (2019). Provides key insights into scaling training corpora, batch sizes, and steps for Transformer pre-training utilized in PEGASUS's scaling experiments.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). Extends abstractive summarization modeling beyond self-supervised pre-training by optimizing summarization policies directly against human feedback and preference models.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Addresses the input sequence length bottleneck in Transformer encoders for summarization by introducing sparse linear attention over long documents.
- Paper: Prefix-Tuning: Optimizing Continuous Prompts for Generation, Xiang Lisa Li et al. (2021). Provides an efficient alternative to full fine-tuning of large sequence-to-sequence models like PEGASUS and BART for downstream summarization tasks via continuous prompt prefixes.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Unifies parameter-efficient transfer methods across large sequence-to-sequence generation benchmarks, providing lighter fine-tuning strategies for pre-trained summarizers.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Presents an advanced, LLM-based reference-free evaluation framework to address the evaluation deficiencies of ROUGE metrics used in PEGASUS.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Augments sequence-to-sequence generators with dynamic retrieval mechanisms to improve factual accuracy in abstractive generation tasks.
- Paper: DiffusER: Discrete Diffusion via Edit-based Reconstruction, Machel Reid et al. (2023). Explores non-autoregressive discrete diffusion generation for iterative editing and abstractive summarization beyond standard autoregressive decoders.
