A trainable document summarizer
J. KupiecJan O. PedersenFrancine R. Chen
Presents a statistical framework that uses a Bayesian classifier to automatically learn feature weights from a corpus and generate concise, indicative document extracts.
Rapidly identifying relevant information within large volumes of full-text documents is a significant operational challenge in digital information retrieval. While author-written abstracts effectively support fast screening, generating coherent narrative summaries automatically remains technically difficult because it demands deep natural language understanding. Document extraction—selecting key sentences directly from the original text—offers a practical and scalable alternative. However, past extraction methods relied on subjective heuristics without a systematic method to combine and weight different selection criteria.
The article evaluates a statistical classification approach to automatic document summarization, demonstrating how to train a probabilistic model to identify and extract key summary sentences using human-written abstracts as benchmarks.
The authors implemented a Bayesian classifier that estimates the probability of each sentence belonging in a summary based on five discrete features: sentence length, fixed indicator phrases, paragraph location, thematic keyword frequency, and uppercase word frequency. The model was trained and evaluated using cross-validation across a dataset of 188 technical document-summary pairs drawn from 21 scientific publications provided by the Engineering Information Company.
The evaluation revealed several key findings regarding summarization performance. First, when generating short summaries equivalent in length to professional abstracts (averaging three sentences), the trained model correctly identified 42% of target sentences, representing a 74% improvement over the standard baseline of taking the opening sentences of a document. Second, the model captured 84% of target summary sentences when permitted to extract 25% of the total document sentences (approximately 20 sentences). Third, structural and location-based heuristics proved most critical; the combination of paragraph position, fixed phrases, and sentence length yielded peak accuracy, whereas content-frequency features showed lower individual utility. Finally, training the classifier on automatically aligned sentence matches achieved 43% precision, showing that the system can be trained effectively without manual data alignment.
These findings indicate that highly effective summarization systems can be built using computationally lightweight statistical methods rather than complex language-generation pipelines. Organizations can implement such trainable extraction tools to enhance search interfaces, accelerate user relevance screening, and reduce document processing costs. Crucially, because the framework learns feature weights directly from training data, it eliminates subjective guesswork and adapts to distinct document collections.
To apply these results, engineering teams should deploy extraction systems that combine location heuristics with phrase detection and integrate supplemental metadata, such as extracted key noun phrases and document titles, to aid rapid browsing. Automated sentence alignment pipelines should be utilized to minimize data preparation costs when training models on new corpora. Future development should evaluate how these models perform across non-technical domains where standard structural cues and fixed phrases may be less prevalent.
Decision-makers should note that the evaluation is based on technical and scientific articles, which exhibit structured formatting that strongly favors location-based heuristics. Additionally, because human summary selection varies significantly—with prior studies showing only 25% overlap among different human judges—precision metrics should be interpreted as strong comparative improvements rather than absolute limits on summary utility.
- Paper: Term-weighting approaches in automatic text retrieval, Gerard Salton et al. (1988). This foundational paper establishes the term-frequency and inverse-document-frequency (TF-IDF) weighting principles that underpin the content-word and keyword scoring features used in the trainable summarizer.
- Paper: A sequential algorithm for training text classifiers, David D. Lewis et al. (1994). It introduces foundational statistical training methodologies for text classification, motivating the transition from heuristic document analysis to machine-learned text processing.
- Paper: The use of MMR, diversity-based reranking for reordering documents and producing summaries, Jaime Carbonell et al. (1998). It extends sentence extraction principles by introducing Maximal Marginal Relevance to penalize redundancy among top-ranked sentences during summary generation.
- Paper: TextRank: Bringing Order into Text, Rada Mihalcea et al. (2004). It advances extractive summarization by modeling inter-sentence relationships using graph centrality algorithms rather than purely local surface features.
- Paper: LexRank: Graph-based Lexical Centrality as Salience in Text Summarization, Günes Erkan et al. (2004). It builds on sentence extraction concepts to formulate graph-based lexical centrality for multi-document extractive summarization.
- Paper: Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics, Chin-Yew Lin et al. (2003). It introduces standard n-gram co-occurrence evaluation metrics (ROUGE precursors) to systematically benchmark summarization systems against human reference abstracts.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). It modernizes both extractive sentence classification and abstractive generation by applying pretrained transformer encoders to document summarization.
- Paper: A Neural Attention Model for Abstractive Sentence Summarization, Alexander M. Rush et al. (2015). It transitions document summarization from feature-based sentence extraction toward data-driven neural attention models capable of generating abstractive summaries.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). It bridges extractive and abstractive paradigms by combining neural generation with a pointer mechanism that selectively copies key source sentences and terms.
- Paper: Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond, Ramesh Nallapati et al. (2016). It incorporates hierarchical sentence-level and word-level representations into sequence-to-sequence neural networks for document summarization.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). It optimizes neural abstractive summarization using deep reinforcement learning objectives to directly improve multi-sentence generation quality.
- Paper: PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, Jingqing Zhang et al. (2020). It introduces a self-supervised pre-training objective tailored specifically for abstractive summarization based on gap-sentence generation.
