Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Rada MihalceaCourtney CorleyCarlo Strapparava

article2006AAAI1,434 citations

Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.

Listen

Modern digital environments rely heavily on short text snippets, such as product descriptions, image captions, and document abstracts. Traditional automated systems typically compare texts by counting exact word matches, but this lexical approach fails when different words share the same meaning (such as "lawyer" and "attorney"). Developing methods that can accurately assess the underlying meaning of short texts is critical for improving information retrieval, text classification, and automated summarization.

The article's main objective is to introduce and evaluate an automated method that measures the semantic similarity between short text segments by combining word-level similarity metrics with word-specificity weighting. The analysis demonstrates that incorporating word meaning significantly improves paraphrase recognition over standard keyword-matching techniques.

To evaluate this approach, the authors developed a scoring formula that matches each open-class word in one text segment to the most semantically related word of the same part of speech in the other text segment, weighting matches by inverse document frequency (a measure of word specificity). The authors tested eight distinct word-level similarity measures—two corpus-based methods derived from large text collections and web search data, and six knowledge-based methods derived from the WordNet semantic hierarchy. These techniques were evaluated in an unsupervised test on the Microsoft Research Paraphrase Corpus, consisting of 1,725 evaluated sentence pairs.

The evaluation yielded several key findings. First, combining all semantic similarity metrics achieved an overall classification accuracy of 70.3%, representing an approximate 13.8% error rate reduction compared to the traditional vector-based cosine baseline (65.4% accuracy). Second, corpus-based pointwise mutual information performed best among individual measures at 69.9% accuracy, while individual knowledge-based measures achieved comparable performance ranging from 69.0% to 69.5%. Third, corpus-based methods generally yielded higher recall (up to 95.2%), whereas knowledge-based methods delivered higher precision (up to 72.4%). Finally, semantic relationships accounted for roughly 20% of all identified word connections, proving that meaning-based connections add substantial value beyond exact word overlap.

These findings indicate that organizations processing short texts can achieve meaningful performance gains by integrating semantic word matching without requiring complex supervised model training. Corpus-based techniques offer the advantage of not requiring manually curated lexical databases, whereas knowledge-based techniques provide finer-grained precision. However, because semantic matching increases similarity scores across related words, it can occasionally misclassify distinct texts with overlapping themes as paraphrases.

Organizations seeking to improve text-matching pipelines should consider adopting hybrid similarity measures, balancing corpus-based methods for broad coverage and knowledge-based methods for precision. For future research and system improvements, the article recommends moving beyond the current unordered word approach to explore structured representations, such as semantic parse trees and predicate logic, to account for sentence grammar and argument roles.

The study's primary limitation is its reliance on a bag-of-words model that ignores grammatical structure and word order, along with restricting word comparisons to identical parts of speech. Confidence in the core conclusion—that semantic metrics statistically outperform standard vector baselines—is high, bounded by the dataset's human annotator agreement level of approximately 83%.

Cover for Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Abstract

This paper presents a method for measuring the semantic similarity of texts, using corpus-based and knowledge-based measures of similarity. Previous work on this problem has focused mainly on either large documents (e.g. text classification, information retrieval) or individual words (e.g. synonymy tests). Given that a large fraction of the information available today, on the Web and elsewhere, consists of short text snippets (e.g. abstracts of scientific documents, image captions, product descriptions), in this paper we focus on measuring the semantic similarity of short texts. Through experiments performed on a paraphrase data set, we show that the semantic similarity method outperforms methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.

Table of Contents

  • Introduction
  • Text Semantic Similarity
  • Semantic Similarity of Words
  • Corpus-based Measures
  • Knowledge-based Measures
  • A Walk-Through Example
  • Evaluation and Results
  • Baselines
  • Results
  • Discussion and Conclusions
  • References

Knowls

  1. Knowl 1 — Text Semantic Similarity Scoring Function

    equation

    Given two text segments T1T_1 and T2T_2, the text semantic similarity score sim(T1,T2)∈[0,1]\text{sim}(T_1, T_2) \in [0, 1] is computed by combining bidirectional maximum word semantic similarities weighted by inverse document frequency:

    sim(T1,T2)=12(∑w∈T1(maxSim(w,T2)⋅idf(w))∑w∈T1idf(w)+∑w∈T2(maxSim(w,T1)⋅idf(w))∑w∈T2idf(w))\text{sim}(T_1, T_2) = \frac{1}{2} \left( \frac{\sum_{w \in T_1} (\text{maxSim}(w, T_2) \cdot \text{idf}(w))}{\sum_{w \in T_1} \text{idf}(w)} + \frac{\sum_{w \in T_2} (\text{maxSim}(w, T_1) \cdot \text{idf}(w))}{\sum_{w \in T_2} \text{idf}(w)} \right)

    where ww denotes an open-class word or cardinal number in the given text segment, idf(w)\text{idf}(w) is the inverse document frequency of word ww derived from a large reference corpus, and maxSim(w,Tj)\text{maxSim}(w, T_j) is the maximum semantic similarity score between ww and any word in segment TjT_j belonging to the same part-of-speech category as ww. A score of 11 signifies semantically identical texts, whereas 00 indicates no semantic overlap.

  2. Knowl 2 — Word Matching and Part-of-Speech Constraints for Text Similarity

    model/method

    The alignment process between words of two text segments T1T_1 and T2T_2 follows specific filtering and matching rules:

    • Word Filtering: All closed-class (function) words are removed; only open-class words (nouns, verbs, adjectives, adverbs) and cardinal numbers participate in semantic matching.
    • Part-of-Speech Restriction: For any word w∈T1w \in T_1, candidate target words in T2T_2 are constrained to share the same part-of-speech (POS) tag as ww.
    • Fallback for Uncovered Word Classes: For word classes where a given semantic similarity metric is not defined (e.g., adjectives or adverbs in taxonomy path-based metrics), an exact lexical match is used as a fallback, assigning maxSim(w,T2)=1.0\text{maxSim}(w, T_2) = 1.0 if w∈T2w \in T_2 and 0.00.0 otherwise.
    • Sense Disambiguation Maximization: For metrics defined over concept/synset pairs rather than raw words, the word-to-word similarity is computed as the maximum concept similarity across all possible sense pairs (c1,c2)(c_1, c_2) of words w1w_1 and w2w_2.
  3. Knowl 3 — Corpus-Based Word Semantic Similarity Metrics

    model/method

    Two corpus-based methods quantify word semantic similarity using statistics extracted from large text corpora:

    1. Pointwise Mutual Information - Information Retrieval (PMI-IR): Measures statistical dependence between words w1w_1 and w2w_2 using search engine query hit counts:

    PMI-IR(w1,w2)=log⁡2hits(w1 NEAR w2)⋅WebSizehits(w1)⋅hits(w2)\text{PMI-IR}(w_1, w_2) = \log_2 \frac{\text{hits}(w_1 \text{ NEAR } w_2) \cdot \text{WebSize}}{\text{hits}(w_1) \cdot \text{hits}(w_2)}

    where hits(w)\text{hits}(w) is the document frequency of word ww, hits(w1 NEAR w2)\text{hits}(w_1 \text{ NEAR } w_2) is the count of documents where w1w_1 and w2w_2 co-occur within a 10-word window, and WebSize=7×1011\text{WebSize} = 7 \times 10^{11} is the estimated size of the indexed Web corpus.

    1. Latent Semantic Analysis (LSA): Applies Singular Value Decomposition (SVD) to a term-by-document matrix TT of a large corpus (such as the British National Corpus), approximating T≈UΣk′VTT \approx U \Sigma_{k'} V^T at reduced dimensionality k′k'. Word and pseudo-document vectors in the reduced LSA space are compared using cosine similarity with tf⋅idf\text{tf} \cdot \text{idf} weighting.

    All raw word similarity scores are normalized to the range [0,1][0, 1] by dividing by the maximum score possible for each metric.

  4. Knowl 4 — Knowledge-Based Word Semantic Similarity Metrics on WordNet

    model/method

    Six knowledge-based concept similarity metrics defined on the WordNet taxonomy are adapted for word-level semantic similarity, each normalized to [0,1][0, 1]:

    • Leacock & Chodorow: Simlch(c1,c2)=−log⁡length(c1,c2)2⋅D\text{Sim}_{\text{lch}}(c_1, c_2) = -\log \frac{\text{length}(c_1, c_2)}{2 \cdot D} where length(c1,c2)\text{length}(c_1, c_2) is the shortest path length between concepts c1c_1 and c2c_2 in the taxonomy, and DD is the maximum taxonomy depth.

    • Lesk: Measures the semantic relatedness between concepts as the extent of word overlap between their dictionary definitions (glosses).

    • Wu & Palmer: Simwup(c1,c2)=2⋅depth(LCS(c1,c2))depth(c1)+depth(c2)\text{Sim}_{\text{wup}}(c_1, c_2) = \frac{2 \cdot \text{depth}(\text{LCS}(c_1, c_2))}{\text{depth}(c_1) + \text{depth}(c_2)} where LCS(c1,c2)\text{LCS}(c_1, c_2) is the Least Common Subsumer of c1c_1 and c2c_2, and depth(c)\text{depth}(c) is the path length from the taxonomic root to concept cc.

    • Resnik: Simres(c1,c2)=IC(LCS(c1,c2))\text{Sim}_{\text{res}}(c_1, c_2) = \text{IC}(\text{LCS}(c_1, c_2)) where IC(c)=−log⁡P(c)\text{IC}(c) = -\log P(c) is the information content of concept cc, with P(c)P(c) being the probability of encountering concept cc in a corpus.

    • Lin: Simlin(c1,c2)=2⋅IC(LCS(c1,c2))IC(c1)+IC(c2)\text{Sim}_{\text{lin}}(c_1, c_2) = \frac{2 \cdot \text{IC}(\text{LCS}(c_1, c_2))}{\text{IC}(c_1) + \text{IC}(c_2)}

    • Jiang & Conrath: Simjnc(c1,c2)=1IC(c1)+IC(c2)−2⋅IC(LCS(c1,c2))\text{Sim}_{\text{jnc}}(c_1, c_2) = \frac{1}{\text{IC}(c_1) + \text{IC}(c_2) - 2 \cdot \text{IC}(\text{LCS}(c_1, c_2))}

  5. Knowl 5 — Unsupervised Paraphrase Identification Setup on MSRP

    experimental setup

    The text semantic similarity metric is evaluated on the task of binary paraphrase identification using the Microsoft Research Paraphrase (MSRP) corpus, comprising 4,076 training pairs and 1,725 test pairs of sentences mined from news sources.

    The experimental setup is entirely unsupervised: no training data is used for tuning. For each test sentence pair (T1,T2)(T_1, T_2), the text semantic similarity score sim(T1,T2)\text{sim}(T_1, T_2) is computed, and the pair is predicted as a paraphrase if and only if sim(T1,T2)≥0.50\text{sim}(T_1, T_2) \ge 0.50. Evaluation metrics comprise Accuracy, Precision, Recall, and F-measure against human judgments (human inter-annotator agreement on the dataset is approximately 83%). Performance is benchmarked against:

    1. A Random baseline that assigns paraphrase labels uniformly at random.
    2. A Vector-based baseline that computes cosine similarity between tf⋅idf\text{tf} \cdot \text{idf}-weighted bag-of-words vectors.
  6. Knowl 6 — Paraphrase Recognition Performance on Microsoft Paraphrase Corpus

    data/table

    Performance of individual and combined text semantic similarity measures on the 1,725 test pairs of the Microsoft Research Paraphrase corpus at an unsupervised classification threshold of 0.500.50:

    Metric Accuracy (%) Precision (%) Recall (%) F-measure (%)
    Semantic similarity (corpus-based)
    PMI-IR 69.9 70.2 95.2 81.0
    LSA 68.4 69.7 95.2 80.5
    Semantic similarity (knowledge-based)
    Jiang Conrath (J C) 69.3 72.2 87.1 79.0
    Leacock Chodorow (L C) 69.5 72.4 87.0 79.0
    Lesk 69.3 72.4 86.6 78.9
    Lin 69.3 71.6 88.7 79.2
    Wu Palmer (W P) 69.0 70.2 92.1 80.0
    Resnik 69.0 69.0 96.4 80.4
    Combined (average of all) 70.3 69.6 97.7 81.3
    Baselines
    Vector-based (tf.idf cosine) 65.4 71.6 79.5 75.3
    Random 51.3 68.3 50.0 57.8

    All semantic similarity metrics achieve a statistically significant improvement over the vector-based cosine baseline (p<0.001p < 0.001, paired tt-test). The combined metric yields an overall accuracy of 70.3% and an F-measure of 81.3%, representing a 13.8% error rate reduction relative to the vector-based baseline (65.4% accuracy). Across metric classes, corpus-based methods achieve higher recall (≥95.2%\ge 95.2\%), whereas knowledge-based methods generally deliver higher precision (up to 72.4%).

  7. Knowl 7 — Pearson Correlation Among Text Semantic Similarity Measures

    data/table

    Pairwise Pearson correlation coefficients between text similarity measures evaluated on the Microsoft Research Paraphrase test set:

    Vect PMI-IR LSA JC LC Lesk Lin WP Resnik
    Vect 1.00 0.84 0.44 0.61 0.63 0.60 0.61 0.50 0.65
    PMI-IR 1.00 0.58 0.67 0.68 0.65 0.67 0.58 0.64
    LSA 1.00 0.42 0.44 0.42 0.43 0.34 0.41
    JC 1.00 0.98 0.97 0.99 0.91 0.45
    LC 1.00 0.98 0.98 0.87 0.46
    Lesk 1.00 0.96 0.86 0.43
    Lin 1.00 0.88 0.44
    WP 1.00 0.34
    Resnik 1.00

    Key patterns shown by the correlation matrix:

    • Strong correlation (≥0.86\ge 0.86) exists across most knowledge-based metrics (Jiang & Conrath, Leacock & Chodorow, Lesk, Lin, Wu & Palmer), reflecting overlapping behavior when contextualized in sentence-level scoring.
    • The Resnik metric exhibits low correlation with other knowledge-based measures (0.340.34--0.460.46) but moderate correlation with PMI-IR (0.640.64), due to its reliance on unnormalized corpus information content.
    • LSA has the weakest correlation with all other measures (0.340.34--0.580.58), indicating that reduced-dimension distributional representations capture distinct relational patterns compared to lexical taxonomies and co-occurrence statistics.
  8. Knowl 8 — Proportion of Lexical Versus Semantic Word Match Links

    empirical result

    In an empirical analysis of the word-level matches identified by the semantic similarity framework across the Microsoft Research Paraphrase dataset:

    • Out of approximately 18,000 total word-to-word similarity pairings formed, approximately 14,500 pairings (~80.6%) represent exact lexical matches (w1=w2w_1 = w_2).
    • Approximately 3,500 pairings (~19.4%) represent non-identical semantic similarity matches between related concepts (e.g., lawyer–attorney, supporters–crowd, court–courthouse).

    This indicates that approximately 20% of the cross-sentence alignment strength that drives paraphrase detection is derived from semantic relatedness beyond surface lexical identity.

  9. Knowl 9 — Bag-of-Words Assumptions and False Positive Vulnerabilities

    limitation

    The proposed text semantic similarity framework is constrained by two major structural limitations:

    1. Neglect of Sentence Structure: Because the scoring function uses a bag-of-words formulation, it ignores syntactic hierarchies, word order, predicate-argument relationships, and directional semantic logic.
    2. Susceptibility to High-Overlap Non-Paraphrases: For non-paraphrase sentence pairs that share high lexical and semantic vocabulary overlap but convey contradictory or distinct meanings (e.g., "The man wasn't on the ice, but trapped in the rapids..." versus "The man was trapped... right at the edge of the falls"), word semantic matching increases the aggregate similarity score above the 0.50 threshold, leading to false positive predictions where strict lexical matching baselines correctly discriminate them.

Coverage note — The step-by-step numerical trace of the walkthrough example on a single sentence pair was omitted as its methodology is fully covered by the scoring function and the full experimental evaluation.

References

  1. 1.Barnard, C., and Callison-Burch, C. 2005. Paraphrasing with bilingual parallel corpora. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics.
  2. 2.Barzilay, R., and Elhadad, N. 2003. Sentence alignment for monolingual comparable corpora. In Proceedings of the Conference on Empirical Methods in Natural Language Processing..
  3. 3.Berry, M. 1992. Large-scale sparse singular value computations. International Journal of Supercomputer Applications 6(1).
  4. 4.Budanitsky, A., and Hirst, G. 2001. Semantic distance in WordNet: An experimental, application-oriented evaluation of five measures. In Proceedings of the NAACL Workshop on WordNet and Other Lexical Resources.
  5. 5.Chklovski, T., and Pantel, P. 2004. Verbocean: Mining the Web for fine-grained semantic verb relations. In Proceedings of Conference on Empirical Methods in Natural Language Processing..
  6. 6.Dagan, I.; Glickman, O.; and Magnini, B. 2005. The PASCAL recognising textual entailment challenge. In Proceedings of the PASCAL Workshop.
  7. 7.Dolan, W.; Quirk, C.; and Brockett, C. 2004. Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources. In Proceedings of the 20th International Conference on Computational Linguistics.
  8. 8.Jiang, J., and Conrath, D. 1997. Semantic similarity based on corpus statistics and lexical taxonomy. In Proceedings of the International Conference on Research in Computational Linguistics.
  9. 9.Landauer, T. K.; Foltz, P.; and Laham, D. 1998. Introduction to latent semantic analysis. Discourse Processes 25.
  10. 10.Lapata, M., and Barzilay, R. 2005. Automatic evaluation of text coherence: Models and representations. In Proceedings of the 19th International Joint Conference on Artificial Intelligence.
  11. 11.Leacock, C., and Chodorow, M. 1998. Combining local context and WordNet sense similarity for word sense identification. In WordNet, An Electronic Lexical Database. The MIT Press.
  12. 12.Lesk, M. 1986. Automatic sense disambiguation using machine readable dictionaries: How to tell a pine cone from an ice cream cone. In Proceedings of the SIGDOC Conference 1986.
  13. 13.Lin, C., and Hovy, E. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of Human Language Technology Conference .
  14. 14.Lin, D., and Pantel, P. 2001. Discovery of inference rules for question answering. Natural Language Engineering 7(3).
  15. 15.Lin, D. 1998. An information-theoretic definition of similarity. In Proceedings of the International Conf. on Machine Learning.
  16. 16.McCarthy, D.; Koeling, R.; Weeds, J.; and Carroll, J. 2004. Finding predominant senses in untagged text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics.
  17. 17.Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics.
  18. 18.Papineni, K. 2001. Why inverse document frequency? In Proceedings of the North American Chapter of the Association for Compuatational Linguistics, 25–32.
  19. 19.Patwardhan, S.; Banerjee, S.; and Pedersen, T. 2003. Using measures of semantic relatedness for word sense disambiguation. In Proceedings of the Fourth International Conference on Intelligent Text Processing and Computational Linguistics.
  20. 20.Resnik, P. 1995. Using information content to evaluate semantic similarity. In Proceedings of the 14th International Joint Conference on Artificial Intelligence.
  21. 21.Rocchio, J. 1971. Relevance feedback in information retrieval. Prentice Hall, Ing. Englewood Cliffs, New Jersey.
  22. 22.Salton, G., and Buckley, C. 1997. Term weighting approaches in automatic text retrieval. In Readings in Information Retrieval. San Francisco, CA: Morgan Kaufmann Publishers.
  23. 23.Salton, G., and Lesk, M. 1971. Computer evaluation of indexing and text processing. Prentice Hall, Ing. Englewood Cliffs, New Jersey. 143–180.
  24. 24.Salton, G.; Singhal, A.; Mitra, M.; and Buckley, C. 1997. Automatic textstructuring and summarization. Information Processing and Management 2(32).
  25. 25.Schutze, H. 1998. Automatic word sense discrimination. Computational Linguistics 24(1):97–124.
  26. 26.Sparck-Jones, K. 1972. A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation 28(1):11–21.
  27. 27.Turney, P. 2001. Mining the web for synonyms: PMI-IR versus LSA on TOEFL. In Proceedings of the Twelfth European Conference on Machine Learning (ECML-2001).
  28. 28.Voorhees, E. 1993. Using WordNet to disambiguate word senses for text retrieval. In Proceedings of the 16th annual international ACM SIGIR conference.
  29. 29.Wu, Z., and Palmer, M. 1994. Verb semantics and lexical selection. In Proceedings of the Annual Meeting of the Association for Computational Linguistics.

Citation

MLA
Mihalcea, R. F., et al. “Corpus-based and Knowledge-based Measures of Text Semantic Similarity”. University of North Texas Digital Library (University of North Texas), 2006, https://digital.library.unt.edu/ark:/67531/metadc30981/.
APA
Mihalcea, R. F., Corley, C. D., & Strapparava, C. (2006). Corpus-based and Knowledge-based Measures of Text Semantic Similarity. University of North Texas Digital Library (University of North Texas). https://digital.library.unt.edu/ark:/67531/metadc30981/
Chicago
Mihalcea, R. F., C. D. Corley, and C. Strapparava. 2006. “Corpus-based and Knowledge-based Measures of Text Semantic Similarity”. University of North Texas Digital Library (University of North Texas). https://digital.library.unt.edu/ark:/67531/metadc30981/.
Harvard
Mihalcea, R.F., Corley, C.D. and Strapparava, C. (2006) “Corpus-based and Knowledge-based Measures of Text Semantic Similarity”, University of North Texas Digital Library (University of North Texas) [Preprint]. Available at: https://digital.library.unt.edu/ark:/67531/metadc30981/.
Vancouver
1. Mihalcea RF, Corley CD, Strapparava C (2006) Corpus-based and Knowledge-based Measures of Text Semantic Similarity. University of North Texas Digital Library (University of North Texas)

BibTeX

@article{mihalcea2006corpus,
  title = {Corpus-based and Knowledge-based Measures of Text Semantic Similarity},
  author = {Mihalcea, Rada F. and Corley, Courtney D. and Strapparava, Carlo},
  year = {2006},
  journal = {University of North Texas Digital Library (University of North Texas)},
  url = {https://digital.library.unt.edu/ark:/67531/metadc30981/}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF