SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents
Ramesh NallapatiFeifei ZhaiBowen Zhou
Introduces SummaRuNNer, a recurrent neural network for extractive document summarization that trains directly on human-written abstractive summaries without sentence-level labels while breaking down sentence selection into interpretable factors like salience, content, and novelty.
Automated document summarization is essential for managing large volumes of textual data in information retrieval and natural language processing. While abstractive summarization generates new text, extractive summarization selects key existing sentences directly from the source text. Extractive techniques remain highly practical because they are computationally less demanding and naturally preserve grammatical and factual correctness. However, many existing deep learning approaches operate as black boxes and require costly, manually labeled sentence datasets for training.
The article introduces and evaluates SummaRuNNer, an interpretable recurrent neural network designed for extractive single-document summarization. The model processes documents sequentially to decide which sentences to include based on clear linguistic factors, and it introduces a training framework capable of learning directly from human-written summaries without requiring sentence-level labels.
The researchers designed a two-layer neural sequence classifier that models text at both the word and sentence levels. Sentence selection decisions explicitly incorporate measurable factors, including information content, document-level salience, redundancy relative to previously chosen sentences, and sentence positioning. The model was trained and evaluated on large-scale datasets, including the CNN and Daily Mail corpora containing over 280,000 news articles, and further tested on the out-of-domain DUC 2002 dataset comprising 567 documents. The authors benchmarked the system against baseline heuristics, traditional graph-based methods, and leading neural extractive and abstractive models.
The findings demonstrate three core results. First, on the Daily Mail benchmark at short summary lengths (75 bytes), SummaRuNNer achieved top-tier performance, outperforming previous deep learning models with significant improvements across standard evaluation metrics (such as a 15% relative improvement in unigram overlap over earlier leading extractive networks). Second, on the combined CNN and Daily Mail benchmark, the extractive model significantly surpassed state-of-the-art abstractive systems, establishing a clear performance advantage over generative approaches. Third, the model's transparent architecture successfully enables visual interpretability, allowing users to inspect exactly how content, salience, novelty, and position contribute to each sentence selection. Additionally, the novel training approach that learns directly from human reference summaries performed competitively, though it trailed direct extractive training by a small margin.
These results show that high-performing summarization can be achieved without the high computational complexity, hallucination risks, or large annotation expenses associated with alternative neural systems. By avoiding the need for dedicated sentence-level manual labeling, organizations can substantially reduce data preparation costs. Furthermore, the ability to decompose and visualize sentence scores directly addresses risk and compliance needs in settings where algorithmic decisions must be explainable.
Organizations deploying automated summarization should consider extractive sequence models as an efficient, low-risk alternative to generative text models. For future technical development, the source suggests combining extractive and abstractive architectures, such as using abstractive models for pre-training or building joint hybrid pipelines where extractive selections feed into text generation modules.
A key operational limitation is domain transferability: when tested on out-of-domain data like the DUC 2002 dataset, the supervised neural model performed below traditional unsupervised graph algorithms, indicating sensitivity to domain shifts. While confidence in the model's performance on news-style corpora is high, decision-makers should exercise caution and conduct domain-specific testing before deploying the system in distinct textual domains.
- Paper: Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond, Ramesh Nallapati et al. (2016). This foundational work adapts attentional sequence-to-sequence RNNs and hierarchical document representations for summarization on the CNN/Daily Mail corpus, establishing the neural summarization framework that SummaRuNNer adapts to the extractive setting.
- Paper: A Neural Attention Model for Abstractive Sentence Summarization, Alexander M. Rush et al. (2015). This paper pioneered the use of neural attention and encoder-decoder networks for text summarization, establishing the core sequence-modeling principles that underpin neural sentence scoring in SummaRuNNer.
- Paper: Document Modeling with Gated Recurrent Neural Network for Sentiment Classification, Duyu Tang et al. (2015). It introduces a hierarchical gated recurrent neural architecture for composing word vectors into sentences and documents, which directly informs SummaRuNNer's two-level RNN document representation.
- Paper: Teaching Machines to Read and Comprehend, Karl Moritz Hermann et al. (2015). This paper introduces the large-scale CNN/Daily Mail corpus and reading comprehension tasks that provided the empirical benchmark used for training and evaluating SummaRuNNer.
- Paper: LexRank: Graph-based Lexical Centrality as Salience in Text Summarization, Günes Erkan et al. (2004). Understanding this classical graph-based lexical centrality baseline is essential for appreciating the extractive sentence-salience modeling that SummaRuNNer reformulates into an end-to-end neural sequence framework.
- Paper: ROUGE: A Package for Automatic Evaluation of Summaries, Chin-Yew Lin (2004). This paper defines the standard ROUGE metric suite required to understand the evaluation methodology and automated reward signals used in SummaRuNNer.
- Paper: A trainable document summarizer, J. Kupiec et al. (1995). This foundational paper establishes the statistical feature-based paradigm (salience, position, and length) for extractive sentence classification that SummaRuNNer explicitly models as interpretable neural components.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). This work advances beyond extractive sentence selection by introducing pointer-generator networks and coverage mechanisms to generate abstractive summaries while directly mitigating repetition and factual errors.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). Building on neural sequence summarization and extractive baselines, this paper introduces a deep reinforcement learning framework combining intra-attention and policy gradients on CNN/Daily Mail.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). This paper generalizes document-level extractive and abstractive summarization by replacing recurrent encoders like SummaRuNNer with pretrained BERT representations.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). This work critiques standard extractive biases demonstrated in earlier benchmarks and explores high-compression extreme summarization using topic-conditioned convolutional architectures.
- Paper: PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, Jingqing Zhang et al. (2020). This study advances neural summarization architectures to large-scale self-supervised pre-training via gap-sentence generation designed specifically for abstractive and extractive tasks.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). This research analyzes faithfulness and factual hallucinations in neural summarization systems, evaluating failure modes that arise when moving beyond extractive models like SummaRuNNer.
