PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

Jingqing ZhangYao ZhaoMohammad SalehPeter J. Liu

article2020ICML2,570 citations

Proposes a gap-sentence pre-training objective specifically designed for abstractive text summarization, setting state-of-the-art performance across 12 diverse datasets while demonstrating remarkable sample efficiency with as few as 1,000 fine-tuning examples.

Listen

The article addresses the challenge of improving abstractive text summarization through pre-training of Transformer models. While general self-supervised pre-training has advanced many NLP tasks, objectives specifically suited to summarization remained unexplored, and evaluations lacked breadth across domains.

This work sets out to develop and test a new pre-training objective called Gap Sentences Generation for encoder-decoder Transformers.

The approach involves pre-training on large corpora by masking important sentences and reconstructing them from the rest of the document. Models were trained on C4 and HugeNews datasets, with ablations performed on smaller models before scaling to 568 million parameters. Performance was assessed on 12 diverse summarization datasets using ROUGE metrics and human evaluations.

Key findings show that the best PEGASUS model achieves state-of-the-art results on all 12 tasks. It also delivers strong performance in low-resource scenarios, exceeding prior benchmarks on six datasets using just 1,000 examples. Human judges rated its summaries as comparable to reference ones on several datasets.

These results indicate that task-aligned pre-training objectives can substantially boost summarization quality and adaptability, particularly when labeled data is limited. Alignment between pre-training and downstream domains further enhances transfer.

Next steps include exploring larger models, handling longer inputs, and applying similar ideas to other generation tasks.

Limitations involve potential data overlap effects, though minimal, and the focus on specific model architectures. Confidence is high given extensive experiments and human validation.

Cover for PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

Abstract

Recent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore there is a lack of systematic evaluation across diverse domains. In this work, we propose pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective. In PEGASUS, important sentences are removed/masked from an input document and are generated together as one output sequence from the remaining sentences, similar to an extractive summary. We evaluated our best PEGASUS model on 12 downstream summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills. Experiments demonstrate it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores. Our model also shows surprising performance on low-resource summarization, surpassing previous state-of-the-art results on 6 datasets with only 1000 examples. Finally we validated our results using human evaluation and show that our model summaries achieve human performance on multiple datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Pre-training Objectives
  • 3.1 Gap Sentences Generation (GSG)
  • 3.2 Masked Language Model (MLM)
  • 4 Pre-training Corpus
  • 5 Downstream Tasks/Datasets
  • 6 Experiments
  • 6.1 Ablations on PEGASUSBASE\text{PEGASUS}_{\text{BASE}}
  • 6.1.1 Pre-training Corpus
  • 6.1.2 Effect of Pre-training Objectives
  • 6.1.3 Effect of Vocabulary
  • 6.2 Larger Model Results
  • 6.3 Zero and Low-Resource Summarization
  • 6.4 Qualitative Observations and Human Evaluation
  • 6.5 Test-set Overlap with Pre-training Corpus
  • 6.6 Additional PEGASUSLARGE\text{PEGASUS}_{\text{LARGE}} Improvements
  • 7 Conclusion
  • 8 Code and Model Checkpoints Release
  • References
  • A Datasets Statistics
  • B Pre-training Steps
  • C PEGASUS Hyper Parameters
  • D Experiment Figures’ Numbers
  • E Low Resource Numbers
  • F Human Evaluation Details
  • G Example of summary with relatively low ROUGE2-F but qualitatively good.
  • H Abstractiveness of Summaries
  • I Example Model Outputs

Knowls

  1. Knowl 1 — Gap Sentences Generation Pre-training Objective

    model/method

    Gap Sentences Generation (GSG) is a self-supervised pre-training objective designed for sequence-to-sequence Transformer models tailored to abstractive summarization. Given an unlabeled document D={x1,x2,…,xn}D = \{x_1, x_2, \dots, x_n\} consisting of nn sentences, GSG selects mm sentences to mask from the input document and concatenates them into a single pseudo-summary target sequence YY for the decoder to generate. In the input document, each selected gap sentence is replaced by a mask token [MASK1][\text{MASK1}].

    The proportion of masked sentences is determined by the gap sentences ratio (GSR), defined as GSR=m/n\text{GSR} = m / n.

    Candidate sentence selection strategies include:

    • Random: mm sentences are chosen uniformly at random without replacement.
    • Lead: The first mm sentences of the document are selected.
    • Principal (Ind-Orig / Ind-Uniq): Sentences are scored independently based on an importance proxy measured by ROUGE1-F1 between each individual sentence xix_i and the remainder of the document: si=rouge(xi,D∖{xi}),∀i∈{1,…,n}s_i = \text{rouge}(x_i, D \setminus \{x_i\}), \quad \forall i \in \{1, \dots, n\} The top mm scoring sentences are chosen. In the "Orig" variant, standard multiset n-gram counting is used; in the "Uniq" variant, n-grams are treated as unique sets to prevent double-counting repeated n-grams.
    • Sequential Principal (Seq-Orig / Seq-Uniq): Sentences are selected greedily and sequentially to maximize the ROUGE1-F1 score between the accumulated selected set S∪{xi}S \cup \{x_i\} and the remaining document D∖(S∪{xi})D \setminus (S \cup \{x_i\}).

    In the final PEGASUS model, independent principal sentence selection with original ROUGE scoring (Ind-Orig) is used. To encourage copying alongside abstractive generation, 20%20\% of the selected gap sentences are randomly kept unchanged in the input rather than replaced by [MASK1][\text{MASK1}].

  2. Knowl 2 — Downstream Summarization Performance of PEGASUS on 12 Diverse Datasets

    data/table

    The PEGASUS models (PEGASUSBASE_{\text{BASE}} with 223M parameters and PEGASUSLARGE_{\text{LARGE}} with 568M parameters) pre-trained with Gap Sentences Generation (GSG) were evaluated across 12 downstream abstractive summarization datasets covering news (XSum, CNN/DailyMail, NEWSROOM, Multi-News, Gigaword), scientific papers (arXiv, PubMed), patents (BIGPATENT), instructions (WikiHow), informal internet stories (Reddit TIFU), emails (AESLC), and legislative bills (BillSum).

    Evaluation results measured by ROUGE-1 (R1), ROUGE-2 (R2), and ROUGE-L (RL) F1 scores against prior state of the art (SOTA) and an un-pretrained Transformer baseline (TransformerBASE_{\text{BASE}}) are shown below:

    Dataset Dataset Size TransformerBASE_{\text{BASE}} PEGASUSBASE_{\text{BASE}} Previous SOTA PEGASUSLARGE_{\text{LARGE}} (C4) PEGASUSLARGE_{\text{LARGE}} (HugeNews)
    XSum 226k 30.83 / 10.83 / 24.41 39.79 / 16.58 / 31.70 45.14 / 22.27 / 37.25 45.20 / 22.06 / 36.99 47.21 / 24.56 / 39.25
    CNN/DailyMail 311k 38.27 / 15.03 / 35.48 41.79 / 18.81 / 38.93 44.16 / 21.28 / 40.90 43.90 / 21.20 / 40.76 44.17 / 21.47 / 41.11
    NEWSROOM 1212k 40.28 / 27.93 / 36.52 42.38 / 30.06 / 38.52 39.91 / 28.38 / 36.87 45.07 / 33.39 / 41.28 45.15 / 33.51 / 41.33
    Multi-News 56k 34.36 / 5.42 / 15.75 42.24 / 13.27 / 21.44 43.47 / 14.89 / 17.41 46.74 / 17.95 / 24.26 47.52 / 18.72 / 24.91
    Gigaword 3995k 35.70 / 16.75 / 32.83 36.91 / 17.66 / 34.08 39.14 / 19.92 / 36.57 38.75 / 19.96 / 36.14 39.12 / 19.86 / 36.24
    WikiHow 168k 32.48 / 10.53 / 23.86 36.58 / 15.64 / 30.01 28.53 / 9.23 / 26.54 43.06 / 19.71 / 34.80 41.35 / 18.51 / 33.42
    Reddit TIFU 42k 15.89 / 1.94 / 12.22 24.36 / 6.09 / 18.75 19.0 / 3.7 / 15.1 26.54 / 8.94 / 21.64 26.63 / 9.01 / 21.60
    BIGPATENT 1341k 42.98 / 20.51 / 31.87 43.55 / 20.43 / 31.80 37.52 / 10.63 / 22.79 53.63 / 33.16 / 42.25 53.41 / 32.89 / 42.07
    arXiv 215k 35.63 / 7.95 / 20.00 34.81 / 10.16 / 22.50 41.59 / 14.26 / 23.55 44.70 / 17.27 / 25.80 44.67 / 17.18 / 25.73
    PubMed 133k 33.94 / 7.43 / 19.02 39.98 / 15.15 / 25.23 40.59 / 15.59 / 23.59 45.49 / 19.90 / 27.69 45.09 / 19.56 / 27.42
    AESLC 18k 15.04 / 7.39 / 14.93 34.85 / 18.94 / 34.10 23.67 / 10.29 / 23.44 37.69 / 21.85 / 36.84 37.40 / 21.22 / 36.45
    BillSum 24k 44.05 / 21.30 / 30.98 51.42 / 29.68 / 37.78 40.80 / 23.83 / 33.73 57.20 / 39.56 / 45.80 57.31 / 40.19 / 45.82

    Pre-training yields substantial gains on smaller datasets: ROUGE-2 F1 scores nearly triple on AESLC and quintuple on Reddit TIFU compared to training without pre-training. On datasets with long reference summaries (BIGPATENT, arXiv, PubMed, Multi-News), targets were truncated to 256 tokens during evaluation.

  3. Knowl 3 — Low-Resource and Zero-Shot Abstractive Summarization with PEGASUS

    empirical result

    PEGASUSLARGE_{\text{LARGE}} pre-trained on HugeNews exhibits high sample efficiency when fine-tuned on low-resource splits (N∈{0,10,100,1000,10000}N \in \{0, 10, 100, 1000, 10000\} examples):

    Dataset 0 examples (Zero-shot) 10 examples 100 examples 1k examples 10k examples Previous SOTA
    R1 / R2 / RL R1 / R2 / RL R1 / R2 / RL R1 / R2 / RL R1 / R2 / RL R1 / R2 / RL
    XSum 19.27 / 3.00 / 12.72 19.39 / 3.45 / 14.02 39.07 / 16.44 / 31.27 41.55 / 18.23 / 33.29 44.71 / 21.20 / 36.31 45.14 / 22.27 / 37.25
    CNN/DailyMail 32.90 / 13.28 / 29.38 37.25 / 15.84 / 33.49 40.28 / 18.21 / 37.03 41.72 / 19.35 / 38.31 42.54 / 20.04 / 39.32 44.16 / 21.28 / 40.90
    NEWSROOM 22.06 / 11.86 / 17.76 29.24 / 17.78 / 24.98 33.63 / 21.81 / 29.64 37.26 / 25.34 / 33.12 39.54 / 27.25 / 35.45 39.91 / 28.38 / 36.87
    Multi-News 36.54 / 10.52 / 18.67 39.79 / 12.56 / 20.06 41.04 / 13.88 / 21.52 44.00 / 15.45 / 22.67 44.70 / 16.57 / 23.43 43.47 / 14.89 / 17.41
    Gigaword 23.39 / 7.59 / 20.20 25.32 / 8.88 / 22.55 29.71 / 12.44 / 27.30 32.95 / 13.90 / 30.10 35.13 / 16.36 / 32.61 38.73 / 19.71 / 35.96
    WikiHow 22.59 / 6.10 / 14.44 23.95 / 6.54 / 15.33 25.24 / 7.52 / 17.79 34.35 / 12.17 / 25.84 37.22 / 14.41 / 29.15 28.53 / 9.23 / 26.54
    Reddit TIFU 14.66 / 3.06 / 10.17 15.36 / 2.91 / 10.76 16.64 / 4.09 / 12.92 23.34 / 6.85 / 18.46 25.47 / 8.18 / 20.33 19.0 / 3.7 / 15.1
    BIGPATENT 25.61 / 6.56 / 17.42 28.87 / 8.30 / 19.71 33.52 / 10.82 / 22.87 36.85 / 12.58 / 24.54 34.81 / 12.39 / 24.13 37.52 / 10.63 / 22.79
    arXiv 28.05 / 6.63 / 17.72 31.38 / 8.16 / 17.97 33.06 / 9.66 / 20.11 39.46 / 12.38 / 22.20 40.24 / 14.04 / 23.11 41.59 / 14.26 / 23.55
    PubMed 28.17 / 7.57 / 17.85 33.31 / 10.58 / 20.05 34.05 / 12.75 / 21.12 40.15 / 15.56 / 24.05 41.75 / 16.74 / 24.80 40.59 / 15.59 / 23.59
    AESLC 10.35 / 3.86 / 9.29 11.97 / 4.91 / 10.84 16.05 / 7.20 / 15.32 28.58 / 15.45 / 28.14 36.47 / 20.85 / 35.53 23.67 / 10.29 / 23.44
    BillSum 41.02 / 17.44 / 25.24 40.48 / 18.49 / 27.27 44.78 / 26.40 / 34.40 46.47 / 30.58 / 37.21 50.81 / 34.49 / 40.96 40.80 / 23.83 / 33.73

    Key results include:

    1. Surpassing SOTA with limited examples: With 100 fine-tuning examples, PEGASUSLARGE_{\text{LARGE}} surpasses prior state-of-the-art ROUGE2-F1 on BIGPATENT, Reddit TIFU, and BillSum. With 1,000 examples, it beats prior SOTA on 6 out of 12 datasets (Multi-News, WikiHow, Reddit TIFU, BIGPATENT, AESLC, and BillSum).
    2. Matching full supervision: On 8 of 12 datasets, fine-tuning PEGASUSLARGE_{\text{LARGE}} with only 100 examples matches or exceeds the performance of an un-pretrained TransformerBASE_{\text{BASE}} trained on the full supervised datasets (ranging from 20k to 200k examples).
    3. Zero-shot performance: On CNN/DailyMail in a zero-shot setting, PEGASUSLARGE_{\text{LARGE}} achieves a ROUGE2-F1 score of 13.28 (compared to 8.27 for zero-shot GPT-2). With 1,000 examples, it achieves 19.35 ROUGE2-F1, outperforming prior language model pre-training on Wikipedia fine-tuned on 3,000 examples (13.1 ROUGE2-F1).
  4. Knowl 4 — Human Evaluation of PEGASUS Summaries

    empirical result

    A human evaluation study was conducted on Amazon Mechanical Turk to evaluate model summaries against human-written reference summaries on XSum, CNN/DailyMail, and Reddit TIFU. Evaluators rated summaries on a 1–5 Likert scale (1 = poor summary, 5 = great summary). Each task was evaluated by 3 independent US workers, taking the median rating per summary. Paired tt-tests evaluated whether ratings were significantly different from human references (p<0.01p < 0.01).

    Experiment / Model XSum mean (pp-value) CNN/DailyMail mean (pp-value) Reddit TIFU mean (pp-value)
    Experiment 1: Pretrain Comparison
    Human-written Reference 3.0 (–) 3.1 (–) 3.2 (–)
    PEGASUSLARGE_{\text{LARGE}} (HugeNews) 3.0 (0.6) 3.6 (0.0001) 3.2 (0.7)
    PEGASUSLARGE_{\text{LARGE}} (C4) 3.1 (0.7) 3.5 (0.009) 3.1 (0.3)
    TransformerBASE_{\text{BASE}} 2.0 (3e-10) 2.9 (0.06) 1.4 (5e-23)
    Experiment 2: Low-Resource Supervision
    Human-written Reference 3.2 (–) 3.2 (–) 3.3 (–)
    PEGASUSLARGE_{\text{LARGE}} (HugeNews) 10 examples 2.8 (0.1) 3.4 (0.007) 2.6 (0.006)
    PEGASUSLARGE_{\text{LARGE}} (HugeNews) 100 examples 3.2 (0.5) 3.4 (0.08) 2.1 (4e-8)
    PEGASUSLARGE_{\text{LARGE}} (HugeNews) 1000 examples 3.4 (0.3) 3.6 (0.07) 2.7 (0.01)
    PEGASUSLARGE_{\text{LARGE}} (HugeNews) full supervision 3.4 (0.3) 3.3 (0.1) 2.8 (0.05)

    Key findings:

    • In the fully supervised setting, PEGASUSLARGE_{\text{LARGE}} pre-trained on HugeNews and C4 produced summaries rated at least as good as (or better than) human reference summaries across all three datasets.
    • Under low supervision (10 to 100 examples), PEGASUSLARGE_{\text{LARGE}} was not measurably worse than human references on XSum and CNN/DailyMail. On Reddit TIFU, human-level performance required full supervision.
  5. Knowl 5 — Model Architectures and Hyperparameters of PEGASUS

    model/method

    PEGASUS uses standard Transformer sequence-to-sequence encoder-decoder architectures. Two main parameter scales are defined:

    • PEGASUSBASE_{\text{BASE}}: Number of encoder and decoder layers L=12L = 12, hidden dimension H=768H = 768, feed-forward layer dimension F=3072F = 3072, attention heads A=12A = 12, total parameter count: 223M223\text{M}. Pre-trained with batch size B=256B = 256 for 500k500\text{k} steps on input length Linput=512L_{\text{input}} = 512 and target length Ltarget=256L_{\text{target}} = 256.
    • PEGASUSLARGE_{\text{LARGE}}: L=16L = 16, H=1024H = 1024, F=4096F = 4096, A=16A = 16, total parameter count: 568M568\text{M}. Pre-trained with batch size B=8192B = 8192 for 500k500\text{k} steps on Linput=512L_{\text{input}} = 512 and Ltarget=256L_{\text{target}} = 256.

    Key architectural and training configurations:

    • Positional Encoding: Sinusoidal positional encodings, which allow generalization to longer input sequence lengths (Linput=1024L_{\text{input}} = 1024) during fine-tuning.
    • Optimization: Adafactor optimizer with square root learning rate decay, learning rate 0.10.1 for pre-training, and dropout rate 0.10.1. Label smoothing of 0.10.1 is applied during fine-tuning.
    • Vocabulary: SentencePiece Unigram tokenizer with a vocabulary size of 96k96\text{k} tokens.
    • Decoding: Greedy decoding during ablation studies; beam search with beam size 88 and length penalty α∈[0.6,0.9]\alpha \in [0.6, 0.9] during downstream fine-tuning evaluations.
  6. Knowl 6 — Sequential Sentence Selection Algorithm for Gap Sentences Generation

    algorithm

    The sequential principal sentence selection strategy greedily chooses mm gap sentences from a document D={x1,…,xn}D = \{x_1, \dots, x_n\} by maximizing the ROUGE1-F1 score between the accumulated selected set and the remaining document at each iteration.

    Input: Document D={x1,x2,…,xn}D = \{x_1, x_2, \dots, x_n\} of nn sentences, number of gap sentences mm
    Output: Set of selected gap sentences SS
    S←∅S \leftarrow \emptyset
    for j←1j \leftarrow 1 to mm do
        for each sentence xi∈D∖Sx_i \in D \setminus S do
            si←rouge(S∪{xi},D∖(S∪{xi}))s_i \leftarrow \text{rouge}(S \cup \{x_i\}, D \setminus (S \cup \{x_i\}))
        end for
        k←arg⁡max⁡i{si}k \leftarrow \arg\max_i \{s_i\}
        S←S∪{xk}S \leftarrow S \cup \{x_k\}
    end for
    return SS

    The function rouge(A,B)\text{rouge}(A, B) calculates the ROUGE1-F1 score between text sequence AA and text sequence BB. In the "Orig" variant, matching n-grams are double-counted as a multiset; in the "Uniq" variant, n-grams are counted as unique sets.

  7. Knowl 7 — Pre-training Objective and Masking Strategy Ablations

    empirical result

    Ablations on PEGASUSBASE_{\text{BASE}} (pre-trained for 500k steps on C4, evaluated with greedy decoding on XSum, CNN/DailyMail, WikiHow, and Reddit TIFU) yielded the following comparisons across objective variants, gap sentences ratios (GSR), and Masked Language Model (MLM) combinations:

    Objective Configuration XSum (R1 / R2 / RL) CNN/DailyMail (R1 / R2 / RL) WikiHow (R1 / R2 / RL) Reddit TIFU (R1 / R2 / RL)
    Random (30% GSR) 39.28 / 16.23 / 31.21 41.80 / 18.91 / 38.88 36.27 / 15.47 / 29.67 24.04 / 6.01 / 18.47
    Lead (30% GSR) 39.22 / 16.12 / 31.09 41.70 / 18.78 / 38.85 35.30 / 14.79 / 28.85 23.48 / 5.78 / 18.00
    Ind-Orig (30% GSR) 39.79 / 16.58 / 31.70 41.79 / 18.81 / 38.93 36.58 / 15.64 / 30.01 24.36 / 6.09 / 18.75
    Ind-Uniq (30% GSR) 39.50 / 16.41 / 31.41 41.79 / 18.83 / 38.94 36.26 / 15.47 / 29.69 24.10 / 5.98 / 18.41
    Seq-Orig (30% GSR) 39.22 / 16.27 / 31.11 41.88 / 18.89 / 39.02 36.39 / 15.57 / 29.74 24.09 / 6.15 / 18.55
    Seq-Uniq (30% GSR) 39.50 / 16.39 / 31.40 41.98 / 19.03 / 39.11 36.69 / 15.61 / 29.95 24.25 / 6.17 / 18.67
    MLM solely 37.22 / 14.48 / 29.62 39.33 / 17.34 / 36.65 32.20 / 13.19 / 27.05 21.00 / 3.96 / 16.27
    MLM Ind-Orig 39.08 / 16.21 / 31.20 41.48 / 18.70 / 38.63 35.99 / 15.29 / 29.57 24.19 / 6.16 / 18.70

    Key findings:

    1. Sentence Selection: Principal sentence selection strategies (Ind-Orig and Seq-Uniq) outperform Random and Lead selection. Lead selection achieves reasonable performance on news datasets due to lead bias, but degrades substantially on non-news datasets.
    2. Gap Sentences Ratio (GSR): Optimal GSR is <50%< 50\%. Specifically, 15%15\% GSR is best on CNN/DailyMail (41.88 / 18.98 / 38.97), 30%30\% on XSum (39.61 / 16.51 / 31.48) and Reddit TIFU (24.05 / 6.05 / 18.55), and 45%45\% on WikiHow (36.39 / 15.46 / 29.85). Setting GSR to 75%75\% degrades downstream ROUGE scores significantly (e.g., Reddit TIFU drops to 21.72 / 4.32 / 16.45).
    3. MLM Co-training: Pre-training with MLM alone performs poorly. Combining MLM with GSG (MLM & Ind-Orig) provides minor gains at early checkpoints (100k–200k steps) but inhibits performance at 500k steps.
  8. Knowl 8 — Pre-training Corpora Comparison and Domain Alignment

    empirical result

    The domain alignment between the pre-training corpus and downstream summarization tasks was evaluated by pre-training PEGASUSBASE_{\text{BASE}} on two large text corpora:

    1. C4: 350M Web pages (750 GB of text).
    2. HugeNews: 1.5B news-like articles (3.8 TB of text) collected from news and news-like domains (2013–2019).

    Downstream evaluation results (ROUGE-1 / ROUGE-2 / ROUGE-L F1 scores):

    Pre-training Corpus XSum CNN/DailyMail WikiHow Reddit TIFU
    C4 39.79 / 16.58 / 31.70 41.79 / 18.81 / 38.93 36.58 / 15.64 / 30.01 24.36 / 6.09 / 18.75
    HugeNews 41.63 / 18.47 / 33.48 42.34 / 19.22 / 39.49 34.93 / 14.67 / 28.63 24.11 / 5.99 / 18.57

    Pre-training on HugeNews substantially improves performance on downstream news tasks (XSum and CNN/DailyMail), whereas pre-training on C4 yields higher performance on non-news informal tasks (WikiHow and Reddit TIFU), demonstrating that pre-training models transfer more effectively when pre-training and downstream domains are closely aligned.

  9. Knowl 9 — PEGASUS_LARGE (mixed, stochastic) Pre-training Configuration

    empirical result

    An extended PEGASUSLARGE_{\text{LARGE}} variant (PEGASUSLARGE_{\text{LARGE}} (mixed, stochastic)) was pre-trained incorporating the following modifications:

    1. Corpora Mixture: Pre-trained on a mixture of C4 and HugeNews, weighted proportionally by their number of examples.
    2. Dynamic Gap Sentences Ratio: The GSR was sampled uniformly between 15%15\% and 45%45\% across training examples.
    3. Stochastic Principal Sentence Selection: Importance scores of sentences were perturbed with 20%20\% uniform noise before selection.
    4. Extended Training: Pre-training was extended to 1.5M1.5\text{M} steps (from 500k500\text{k} steps) to accommodate slower convergence.
    5. Tokenizer Update: The SentencePiece tokenizer was updated to explicitly encode newline characters.

    Downstream performance (ROUGE-1 / ROUGE-2 / ROUGE-L F1 scores):

    Dataset ROUGE-1 ROUGE-2 ROUGE-L
    XSum 47.60 24.83 39.64
    CNN/DailyMail 44.16 21.56 41.30
    NEWSROOM 45.98 34.20 42.18
    Multi-News 47.65 18.75 24.95
    Gigaword 39.65 20.47 36.76
    WikiHow 46.39 22.12 38.41
    Reddit TIFU 27.99 9.81 22.94
    BIGPATENT (cased) 52.29 33.08 41.66
    arXiv 44.21 16.95 25.67
    PubMed 45.97 20.15 28.25
    AESLC 37.68 21.25 36.51
    BillSum 59.67 41.58 47.59

    This configuration achieves improved performance across nearly all downstream abstractive summarization benchmarks.

  10. Knowl 10 — Pre-training Data Overlap and Memorization Analysis

    empirical result

    To evaluate whether performance on downstream datasets was influenced by memorization of pre-training web documents, overlap was quantified using ROUGE-2 recall between downstream test set targets and C4 pre-training documents (calculated as common 2-grams/test set target 2-grams\text{common 2-grams} / \text{test set target 2-grams}).

    At similarity thresholds of 1.01.0 and 0.80.8:

    • CNN/DailyMail, Reddit TIFU, and WikiHow exhibited minimal overlap (<5%< 5\% of test examples matched pre-training data at similarity >0.8> 0.8).
    • XSum had an overlap rate between 15%15\% and 20%20\%.
    • Removing all overlapping test examples (at similarity >0.8> 0.8 or =1.0= 1.0) changed downstream ROUGE scores by less than 1%1\%.
    • Manual inspection of test examples with a similarity score of 1.01.0 showed that model decodes differed substantially in structure and wording from reference summaries, indicating that downstream performance is not driven by verbatim memorization.

Coverage note — Omitted qualitative summary samples across various ROUGE brackets provided in Appendix I, as they serve as individual illustrative examples rather than standalone findings.

References

  1. 1.Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014.
  2. 2.Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pp. 615–621, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-2097. URL https://www.aclweb.org/anthology/N18-2097.
  3. 3.Dai, A. M. and Le, Q. V. Semi-supervised sequence learning. In Cortes, C., Lawrence, N. D., Lee, D. D., Sugiyama, M., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 28, pp. 3079–3087. Curran Associates, Inc., 2015. URL http://papers.nips.cc/paper/5949-semi-supervised-sequence-learning.pdf.
  4. 4.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://www.aclweb.org/anthology/N19-1423.
  5. 5.Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H.-W. Unified language model pre-training for natural language understanding and generation. In 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019.
  6. 6.Fabbri, A., Li, I., She, T., Li, S., and Radev, D. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1074–1084, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1102. URL https://www.aclweb.org/anthology/P19-1102.
  7. 7.Goodman, S., Lan, Z., and Soricut, R. Multi-stage pretraining for abstractive summarization, 2019.
  8. 8.Graff, D., Kong, J., Chen, K., and Maeda, K. English gigaword. Linguistic Data Consortium, Philadelphia, 4 (1):34, 2003.
  9. 9.Grusky, M., Naaman, M., and Artzi, Y. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 2018. doi: 10.18653/v1/n18-1065. URL http://dx.doi.org/10.18653/v1/n18-1065.
  10. 10.Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P. Teaching machines to read and comprehend. In Advances in neural information processing systems, pp. 1693–1701, 2015.
  11. 11.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Comput., 9(8):1735–1780, November 1997. ISSN 0899-7667. doi: 10.1162/neco.1997.9.8.1735. URL http://dx.doi.org/10.1162/neco.1997.9.8.1735.
  12. 12.Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O. SpanBERT: Improving pre-training by representing and predicting spans. arXiv preprint arXiv:1907.10529, 2019.
  13. 13.Khandelwal, U., Clark, K., Jurafsky, D., and Kaiser, L. Sample efficient text summarization using a single pretrained transformer. arXiv preprint arXiv:1905.08836, 2019.
  14. 14.Kim, B., Kim, H., and Kim, G. Abstractive summarization of Reddit posts with multi-level memory networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2519–2531, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1260. URL https://www.aclweb.org/anthology/N19-1260.
  15. 15.Klimt, B. and Yang, Y. The enron corpus: A new dataset for email classification research. In Proceedings of the 15th European Conference on Machine Learning, ECML’04, pp. 217–226, Berlin, Heidelberg, 2004. Springer-Verlag. ISBN 3-540-23105-6, 978-3-540-23105-9. doi: 10.1007/978-3-540-30115-8 22. URL https://doi.org/10.1007/978-3-540-30115-8_22.
  16. 16.Kornilova, A. and Eidelman, V. BillSum: A corpus for automatic summarization of US legislation. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pp. 48–56, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-5406. URL https://www.aclweb.org/anthology/D19-5406.
  17. 17.Koupaee, M. and Wang, W. Y. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305, 2018.
  18. 18.Kryscinski, W., Keskar, N. S., McCann, B., Xiong, C., and Socher, R. Neural text summarization: A critical evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 540–551, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1051. URL https://www.aclweb.org/anthology/D19-1051.
  19. 19.Kudo, T. Subword regularization: Improving neural network translation models with multiple subword candidates. arXiv preprint arXiv:1804.10959, 2018.
  20. 20.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019.
  21. 21.Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/W04-1013.
  22. 22.Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. Generating wikipedia by summarizing long sequences. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Hyg0vbWC-.
  23. 23.Nallapati, R., Zhou, B., dos Santos, C., Gulçehre, Ç ., and Xiang, B. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pp. 280–290, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/K16-1028. URL https://www.aclweb.org/anthology/K16-1028.
  24. 24.Nallapati, R., Zhai, F., and Zhou, B. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI’17, pp. 3075–3081. AAAI Press, 2017. URL http://dl.acm.org/citation.cfm?id=3298483.3298681.
  25. 25.Narayan, S., Cohen, S. B., and Lapata, M. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 1797–1807, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1206. URL https://www.aclweb.org/anthology/D18-1206.
  26. 26.Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017.
  27. 27.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. URL https://s3-us-west-2.amazonaws.com/openai-assets/researchcovers/languageunsupervised/language_understanding_paper.pdf, 2018a.
  28. 28.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2018b. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf.
  29. 29.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer, 2019.
  30. 30.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016. doi: 10.18653/v1/d16-1264. URL http://dx.doi.org/10.18653/v1/D16-1264.
  31. 31.Ramachandran, P., Liu, P., and Le, Q. Unsupervised pretraining for sequence to sequence learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 383–391, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1039. URL https://www.aclweb.org/anthology/D17-1039.
  32. 32.Rothe, S., Narayan, S., and Severyn, A. Leveraging pretrained checkpoints for sequence generation tasks. arXiv preprint arXiv:1907.12461, 2019.
  33. 33.Rush, A. M., Chopra, S., and Weston, J. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 379–389, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1044. URL https://www.aclweb.org/anthology/D15-1044.
  34. 34.See, A., Liu, P. J., and Manning, C. D. Get to the point: Summarization with pointer-generator networks. CoRR, abs/1704.04368, 2017. URL http://arxiv.org/abs/1704.04368.
  35. 35.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1715–1725, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1162. URL https://www.aclweb.org/anthology/P16-1162.
  36. 36.Sharma, E., Li, C., and Wang, L. BIGPATENT: A largescale dataset for abstractive and coherent summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2204–2213, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1212. URL https://www.aclweb.org/anthology/P19-1212.
  37. 37.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. arXiv preprint arXiv:1804.04235, 2018.
  38. 38.Shi, T., Wang, P., and Reddy, C. K. LeafNATS: An opensource toolkit and live demo system for neural abstractive text summarization. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pp. 66–71, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-4012. URL https://www.aclweb.org/anthology/N19-4012.
  39. 39.Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y. Mass: Masked sequence to sequence pre-training for language generation. In International Conference on Machine Learning, pp. 5926–5936, 2019.
  40. 40.Subramanian, S., Li, R., Pilault, J., and Pal, C. On extractive and abstractive neural document summarization with transformer language models. arXiv preprint arXiv:1909.03186, 2019.
  41. 41.Sutskever, I., Vinyals, O., and Le, Q. V. Sequence to sequence learning with neural networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2, NIPS’14, pp. 3104–3112, Cambridge, MA, USA, 2014. MIT Press. URL http://dl.acm.org/citation.cfm?id=2969033.2969173.
  42. 42.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  43. 43.Völske, M., Potthast, M., Syed, S., and Stein, B. TL;DR: Mining Reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pp. 59–63, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-4508. URL https://www.aclweb.org/anthology/W17-4508.
  44. 44.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. Glue: A multi-task benchmark and analysis platform for natural language understanding. Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2018. doi: 10.18653/v1/w18-5446. URL http://dx.doi.org/10.18653/v1/w18-5446.
  45. 45.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. arXiv preprint arXiv:1908.04319, 2019.
  46. 46.Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  47. 47.Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, pp. 5754–5764, 2019. URL http://papers.nips.cc/paper/8812-xlnet-generalized-autoregressive-pretraining-for-language-understanding.pdf.
  48. 48.Zhang, R. and Tetreault, J. This email could save your life: Introducing the task of email subject line generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 446–456, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1043. URL https://www.aclweb.org/anthology/P19-1043.
  49. 49.Zhong, M., Liu, P., Wang, D., Qiu, X., and Huang, X. Searching for effective neural extractive summarization: What works and whats next. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. doi: 10.18653/v1/p19-1100. URL http://dx.doi.org/10.18653/v1/p19-1100.

Citation

MLA
Zhang, J., et al. “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization”. International Conference on Machine Learning, vol. 119, 2020, pp. 11328–39, https://proceedings.mlr.press/v119/zhang20ae.html.
APA
Zhang, J., Zhao, Y., Saleh, M., & Liu, P. (2020). PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. International Conference on Machine Learning, 119, 11328–11339. https://proceedings.mlr.press/v119/zhang20ae.html
Chicago
Zhang, J., Y. Zhao, M. Saleh, and P. Liu. 2020. “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization”. International Conference on Machine Learning 119: 11328–39. https://proceedings.mlr.press/v119/zhang20ae.html.
Harvard
Zhang, J. et al. (2020) “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization”, International Conference on Machine Learning. PMLR, pp. 11328–11339. Available at: https://proceedings.mlr.press/v119/zhang20ae.html.
Vancouver
1. Zhang J, Zhao Y, Saleh M, Liu P (2020) PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. In: International Conference on Machine Learning. PMLR, pp 11328–11339

BibTeX

@InProceedings{pmlr-v119-zhang20ae,
  title = 	 {{PEGASUS}: Pre-training with Extracted Gap-sentences for Abstractive Summarization},
  author =       {Zhang, Jingqing and Zhao, Yao and Saleh, Mohammad and Liu, Peter},
  booktitle = 	 {Proceedings of the 37th International Conference on Machine Learning},
  pages = 	 {11328--11339},
  year = 	 {2020},
  editor = 	 {III, Hal Daumé and Singh, Aarti},
  volume = 	 {119},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--18 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v119/zhang20ae/zhang20ae.pdf},
  url = 	 {https://proceedings.mlr.press/v119/zhang20ae.html},
  abstract = 	 {Recent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore there is a lack of systematic evaluation across diverse domains. In this work, we propose pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective. In PEGASUS, important sentences are removed/masked from an input document and are generated together as one output sequence from the remaining sentences, similar to an extractive summary. We evaluated our best PEGASUS model on 12 downstream summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills. Experiments demonstrate it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores. Our model also shows surprising performance on low-resource summarization, surpassing previous state-of-the-art results on 6 datasets with only 1000 examples. Finally we validated our results using human evaluation and show that our model summaries achieve human performance on multiple datasets.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/