A trainable document summarizer

J. KupiecJan O. PedersenFrancine R. Chen

article1995SIGIR1,585 citations

Presents a statistical framework that uses a Bayesian classifier to automatically learn feature weights from a corpus and generate concise, indicative document extracts.

Listen

Rapidly identifying relevant information within large volumes of full-text documents is a significant operational challenge in digital information retrieval. While author-written abstracts effectively support fast screening, generating coherent narrative summaries automatically remains technically difficult because it demands deep natural language understanding. Document extraction—selecting key sentences directly from the original text—offers a practical and scalable alternative. However, past extraction methods relied on subjective heuristics without a systematic method to combine and weight different selection criteria.

The article evaluates a statistical classification approach to automatic document summarization, demonstrating how to train a probabilistic model to identify and extract key summary sentences using human-written abstracts as benchmarks.

The authors implemented a Bayesian classifier that estimates the probability of each sentence belonging in a summary based on five discrete features: sentence length, fixed indicator phrases, paragraph location, thematic keyword frequency, and uppercase word frequency. The model was trained and evaluated using cross-validation across a dataset of 188 technical document-summary pairs drawn from 21 scientific publications provided by the Engineering Information Company.

The evaluation revealed several key findings regarding summarization performance. First, when generating short summaries equivalent in length to professional abstracts (averaging three sentences), the trained model correctly identified 42% of target sentences, representing a 74% improvement over the standard baseline of taking the opening sentences of a document. Second, the model captured 84% of target summary sentences when permitted to extract 25% of the total document sentences (approximately 20 sentences). Third, structural and location-based heuristics proved most critical; the combination of paragraph position, fixed phrases, and sentence length yielded peak accuracy, whereas content-frequency features showed lower individual utility. Finally, training the classifier on automatically aligned sentence matches achieved 43% precision, showing that the system can be trained effectively without manual data alignment.

These findings indicate that highly effective summarization systems can be built using computationally lightweight statistical methods rather than complex language-generation pipelines. Organizations can implement such trainable extraction tools to enhance search interfaces, accelerate user relevance screening, and reduce document processing costs. Crucially, because the framework learns feature weights directly from training data, it eliminates subjective guesswork and adapts to distinct document collections.

To apply these results, engineering teams should deploy extraction systems that combine location heuristics with phrase detection and integrate supplemental metadata, such as extracted key noun phrases and document titles, to aid rapid browsing. Automated sentence alignment pipelines should be utilized to minimize data preparation costs when training models on new corpora. Future development should evaluate how these models perform across non-technical domains where standard structural cues and fixed phrases may be less prevalent.

Decision-makers should note that the evaluation is based on technical and scientific articles, which exhibit structured formatting that strongly favors location-based heuristics. Additionally, because human summary selection varies significantly—with prior studies showing only 25% overlap among different human judges—precision metrics should be interpreted as strong comparative improvements rather than absolute limits on summary utility.

  • Paper: Term-weighting approaches in automatic text retrieval, Gerard Salton et al. (1988). This foundational paper establishes the term-frequency and inverse-document-frequency (TF-IDF) weighting principles that underpin the content-word and keyword scoring features used in the trainable summarizer.
  • Paper: A sequential algorithm for training text classifiers, David D. Lewis et al. (1994). It introduces foundational statistical training methodologies for text classification, motivating the transition from heuristic document analysis to machine-learned text processing.
Cover for A trainable document summarizer

Abstract

• To summarize is to reduce in complexity, and hence in length, while retaining some of the essential qualities of the original.

• This paper focusses on document extracts, a particular kind of computed document summary.

• Document extracts consisting of roughly 20% of the original can be as informative as the full text of a document, which suggests that even shorter extracts may be useful indicative summaries.

• The trends in our results are in agreement with those of Edmundson who used a subjectively weighted combination of features as opposed to training the feature weights using a corpus.

• We have developed a trainable summarization program that is grounded in a sound statistical framework.

Table of Contents

  • 1 Introduction
  • 2 A Trainable Summarizer
  • 2.1 Features
  • 2.2 Classifier
  • 3 The Corpus
  • 3.1 Sentence Matching
  • 4 Evaluation
  • 5 Discussion
  • 6 Implementation Issues
  • 7 Conclusions
  • 8 Acknowledgments
  • References
  • 9 Appendix
  • 9.0.1 Direct Match
  • 9.0.2 Direct Join
  • 9.0.3 Incomplete Matches

Knowls

  1. Knowl 1 — Naive Bayes Classification Framework for Sentence Extraction

    model/method

    Extractive document summarization is formulated as a statistical classification problem where each sentence ss in a document is assigned a probability of being included in a summary SS given kk discrete features F1,F2,…,FkF_1, F_2, \dots, F_k. Assuming conditional independence among features, Bayes' rule yields the posterior probability:

    P(s∈S∣F1,F2,…,Fk)=P(s∈S)∏j=1kP(Fj∣s∈S)∏j=1kP(Fj)P(s \in S \mid F_1, F_2, \dots, F_k) = \frac{P(s \in S) \prod_{j=1}^k P(F_j \mid s \in S)}{\prod_{j=1}^k P(F_j)}

    where:

    • ss denotes an individual sentence from the source document.
    • SS denotes the target document extract or summary.
    • P(s∈S)P(s \in S) is the constant prior probability that any given sentence is selected for a summary.
    • P(Fj∣s∈S)P(F_j \mid s \in S) is the conditional probability of observing feature FjF_j given that the sentence is part of the summary.
    • P(Fj)P(F_j) is the marginal probability of observing feature FjF_j.

    Because all features FjF_j are discrete, P(Fj∣s∈S)P(F_j \mid s \in S) and P(Fj)P(F_j) are estimated directly from training data via frequency counts. Sentences in an unseen document are scored according to this posterior probability, ranked in descending order, and the top-scoring sentences are selected to form the extract.

  2. Knowl 2 — Discrete Feature Set for Sentence Extraction

    model/method

    The trainable summarizer employs five discrete sentence-level features:

    1. Sentence Length Cut-off Feature: A binary feature evaluated against a fixed length threshold (e.g., 5 words). It is true for sentences with word counts strictly greater than the threshold and false otherwise, filtering out short headers and fragments.
    2. Fixed-Phrase Feature: A binary feature indicating whether a sentence contains any of 26 predefined indicator phrases (such as "this letter..." or "in conclusion...") or immediately follows a section heading containing keywords such as "conclusions", "results", "summary", or "discussion".
    3. Paragraph Feature: A discrete categorical feature defined over the first 10 paragraphs and the last 5 paragraphs of a document. It categorizes sentences by their relative position within each paragraph: paragraph-initial, paragraph-final (in paragraphs of length ≥2\ge 2 sentences), or paragraph-medial (in paragraphs of length >2> 2 sentences).
    4. Thematic Word Feature: A binary feature based on the most frequent content words in the document (thematic words, with frequency ties broken by word length). Sentences receive a score as a function of the frequency of these thematic words; the feature evaluates to true if the sentence ranks among the top-scoring set of sentences in the document.
    5. Uppercase Word Feature: A binary feature that scores proper names and acronym expansions (e.g., "... by the ASTM (American Society for Testing and Materials)"). It identifies capitalized content words that are not sentence-initial, appear multiple times in the text, and are not abbreviated units of measurement (such as F, C, or Kg). Sentences containing the first occurrence of such an uppercase word receive double the score weight of subsequent occurrences; the feature evaluates to true for sentences among the highest scoring set.
  3. Knowl 3 — Individual and Cumulative Feature Performance in Sentence Extraction

    data/table

    In a leave-one-journal-out cross-validation across 498 matchable sentences from 188 technical articles, individual features and their cumulative combinations (added in descending order of individual accuracy) yielded the following sentence-level precision counts when outputting the same number of sentences as in human-authored summaries:

    Feature Individual Sents Correct Cumulative Sents Correct
    Paragraph 163 (33%) 163 (33%)
    Fixed Phrases 145 (29%) 209 (42%)
    Length Cut-off 121 (24%) 217 (44%)
    Thematic Word 101 (20%) 209 (42%)
    Uppercase Word 100 (20%) 211 (42%)

    Paragraph position is the single most informative feature (33%), followed by Fixed Phrases (29%). The optimal combination of features is Paragraph + Fixed Phrases + Length Cut-off, achieving 44% precision (217 of 498 sentences). Adding frequency-keyword features (Thematic Word and Uppercase Word) slightly lowers cumulative precision to 42% (211 of 498 sentences), because keyword frequency features select sentences evenly throughout the entire document, whereas indicative summary sentences are strongly concentrated at the beginnings and ends of documents.

  4. Knowl 4 — Extractive Summarization Precision across Summary Sizes and Baselines

    empirical result

    When evaluated on a corpus of 188 scientific and technical documents (averaging 86 sentences per document, with professional summaries averaging 3 sentences):

    • Baseline Comparison: The default baseline of selecting sentences sequentially from the beginning of the text (equivalent to using the Sentence Length Cut-off feature alone) correctly identifies 121 sentences (24% precision). Using the full discrete Bayesian classifier achieves 211 correct sentences (42% precision), representing a 74% relative improvement over the lead-text baseline.
    • Stringent Replication Rate: When evaluated against all 568 human summary sentences (where multi-sentence joins require all participating original sentences to be extracted), the summarizer replicates 201 sentences (35% precision), approaching the 83% theoretical upper bound imposed by non-verbatim abstracting (451 direct matches + 19 direct joins out of 568 total sentences).
    • Extract Size vs. Precision: Generating an extract corresponding to 25% of the original document length (approximately 20 sentences for an 86-sentence text) recovers 84% of the sentences selected in human-authored summaries.
  5. Knowl 5 — Taxonomy of Summary-to-Document Sentence Correspondences

    definition

    To train and evaluate extractive summarizers against non-verbatim human-written abstracts, correspondences between manual summary sentences and original document sentences are categorized into five types:

    1. Direct Sentence Match: A summary sentence is extracted verbatim from a single original sentence, or with minor stylistic modifications that fully preserve the underlying content.
    2. Direct Join: A summary sentence is constructed by combining two or more original document sentences, preserving the combined content of all component sentences.
    3. Unmatchable Sentence: A summary sentence is written based on a general reading of the source document without corresponding to any specific original sentence.
    4. Incomplete Single Sentence: An original sentence shares partial topical overlap with a summary sentence, but the summary sentence either does not preserve the full content or includes additional information not found in that source sentence.
    5. Incomplete Join: A combination of original sentences that only partially covers the content of a summary sentence or includes ungrounded external information.
  6. Knowl 6 — Trainability from Automated Sentence Alignment without Manual Correction

    empirical result

    A Bayesian sentence extraction summarizer trained solely on document-summary sentence correspondences generated by an automated sentence alignment program (with zero manual alignment or cleaning) correctly identified 216 sentences (43% precision) on 498 matchable test sentences under cross-validation. This matches the 211 sentences (42% precision) achieved when training on manually verified and corrected sentence alignments, demonstrating that the statistical summarization model can be trained end-to-end on uncurated corpora of document-summary pairs without requiring human alignment intervention.

  7. Knowl 7 — Document Preprocessing and Key Phrase Presentation Pipeline

    model/method

    The end-to-end document summarization system incorporates two complementary non-extraction components for robust real-world presentation:

    1. Boundary and Title Identification: Because document bodies are frequently preceded by metadata (author names, affiliations, publication dates), a specialized feature set using capitalization patterns, sentence punctuation, numerical tokens, and sentence/paragraph lengths identifies the true start of the main text body and isolates the title sentence. If an explicit author abstract header is detected, its text is directly used as the summary without running sentence extraction.
    2. Key Phrase Extraction: To complement extracted sentences, key phrases consisting of adjacent words (primarily noun phrases of length ≥2\ge 2) are extracted by filtering sentences through a stop list of articles, prepositions, common adverbs, and auxiliary verbs. Candidate phrases are ranked by corpus and document frequency and presented alongside the extracted sentences in reading order.

Coverage note — None was omitted; all core mathematical formulations, feature descriptions, data distributions, experimental findings, alignment comparisons, and presentation pipeline components are fully covered.

References

  1. 1.P. B. Baxendale. Man-made index for technical literature – an experiment. IBM J. Res. Develop., 2(4):354–361, 1958.
  2. 2.F.R. Chen and M.M. Withgott. The use of emphasis to automatically summarize a spoken discourse. In Proceedings of the IEEE Intl. Conf. on Acoust., Speech and Signal Proc., volume 1, pages 229–232, March 1992.
  3. 3.G. DeJong. An overview of the FRUMP system. In W.G. Lehnert and M. H. Ringle, editors, Strategies for Natural Language Parsing, pages 149–176, 1982.
  4. 4.H. P. Edmundson. New methods in automatic abstracting. Journal of the ACM, 16(2):264–285, April 1969.
  5. 5.P.S. Jacobs and L. F. Rau. Scisor: Extracting information from on-line news. Communications of the ACM, 33(11):88–97, 1990.
  6. 6.K. Sparck Jones. Discourse modelling for automatic summarizing. Technical Report 290, Computer Laboratory, University of Cambridge, 1993.
  7. 7.H.P. Luhn. The automatic creation of literature abstracts. IBM J. Res. Develop., 2:159–165, 1959.
  8. 8.S. Miike, E. Itoh, K. Ono, and K. Sumita. A full-text retrieval system with a dynamic abstract generation function. In W. Bruce Croft and C.J. van Rijsbergen, editors, Proceedings of Seventeenth Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 152–161, July 1994.
  9. 9.A. H. Morris, G. M. Kasper, and D. A. Adams. The effects and limitations of automated text condensing on reading comprehension performance. Information Systems Research, pages 17–35, March 1992.
  10. 10.C. D. Paice. Constructing literature abstracts by computer: Techniques and prospects. Information Processing and Management, 26:171–186, 1990.
  11. 11.C. D. Paice and P. A. Jones. The identification of important concepts in highly structured technical papers. In R. Korfhage, E. Rasmussen, and P. Willett, editors, Proceedings of Sixteenth Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 69–78. ACM Press, June 1993.
  12. 12.G. J. Rath, A. Resnick, and T. R. Savage. The formation of abstracts by the selection of sentences. American Documentation, 12(2):139–143, April 1961.
  13. 13.U. Reimer and U. Hahn. Text condensation as knowledge base abstraction. In IEEE Conf. on AI Applications, pages 338–344, 1988.
  14. 14.G. Salton, J. Allan, and C. Buckley. Approaches to passage retrieval in full text information systems. In Proceedings of SIGIR '93, pages 49–58, June 1993.
  15. 15.G. Salton, J. Allan, C. Buckley, and A. Singhal. Automatic analysis, theme generation, and summarization of machine-readable texts. Science, 264(3):1421–1426, June 1994.
  16. 16.C. Schwarz. Content based text handling. Information Processing & Management, 26(2):219–226, 1990.
  17. 17.E. F. Skorokhod'ko. Adaptive method of automatic abstracting and indexing. In IFIP Congress, Ljubljana, Yugoslavia 71, pages 1179–1182. North Holland, 1972.
  18. 18.J. I. Tait. Generating summaries using a script-based language analyzer. In L. Steels and J.A. Campbell, editors, Progress in Artificial Intelligence, pages 312–318. Ellis Horwood, 1985.
  19. 19.L. C. Tong and S. L. Tan. A statistical approach to automatic text extraction. Asian Library Journal.

Citation

MLA
Kupiec, J., et al. “A Trainable Document Summarizer”. Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '95, 1995, pp. 68–73, https://doi.org/10.1145/215206.215333.
APA
Kupiec, J., Pedersen, J., & Chen, F. (1995). A trainable document summarizer. Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '95, 68–73. https://doi.org/10.1145/215206.215333
Chicago
Kupiec, J., J. Pedersen, and F. Chen. 1995. “A Trainable Document Summarizer”. Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '95, 68–73. https://doi.org/10.1145/215206.215333.
Harvard
Kupiec, J., Pedersen, J. and Chen, F. (1995) “A trainable document summarizer”, Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '95. ACM Press, pp. 68–73. Available at: https://doi.org/10.1145/215206.215333.
Vancouver
1. Kupiec J, Pedersen J, Chen F (1995) A trainable document summarizer. In: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '95. ACM Press, pp 68–73

BibTeX

@inproceedings{Kupiec_1995, series={SIGIR ’95}, title={A trainable document summarizer}, url={http://dx.doi.org/10.1145/215206.215333}, DOI={10.1145/215206.215333}, booktitle={Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval  - SIGIR ’95}, publisher={ACM Press}, author={Kupiec, Julian and Pedersen, Jan and Chen, Francine}, year={1995}, pages={68–73}, collection={SIGIR ’95} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF