Corpus-based and Knowledge-based Measures of Text Semantic Similarity
Rada MihalceaCourtney CorleyCarlo Strapparava
Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.
Modern digital environments rely heavily on short text snippets, such as product descriptions, image captions, and document abstracts. Traditional automated systems typically compare texts by counting exact word matches, but this lexical approach fails when different words share the same meaning (such as "lawyer" and "attorney"). Developing methods that can accurately assess the underlying meaning of short texts is critical for improving information retrieval, text classification, and automated summarization.
The article's main objective is to introduce and evaluate an automated method that measures the semantic similarity between short text segments by combining word-level similarity metrics with word-specificity weighting. The analysis demonstrates that incorporating word meaning significantly improves paraphrase recognition over standard keyword-matching techniques.
To evaluate this approach, the authors developed a scoring formula that matches each open-class word in one text segment to the most semantically related word of the same part of speech in the other text segment, weighting matches by inverse document frequency (a measure of word specificity). The authors tested eight distinct word-level similarity measures—two corpus-based methods derived from large text collections and web search data, and six knowledge-based methods derived from the WordNet semantic hierarchy. These techniques were evaluated in an unsupervised test on the Microsoft Research Paraphrase Corpus, consisting of 1,725 evaluated sentence pairs.
The evaluation yielded several key findings. First, combining all semantic similarity metrics achieved an overall classification accuracy of 70.3%, representing an approximate 13.8% error rate reduction compared to the traditional vector-based cosine baseline (65.4% accuracy). Second, corpus-based pointwise mutual information performed best among individual measures at 69.9% accuracy, while individual knowledge-based measures achieved comparable performance ranging from 69.0% to 69.5%. Third, corpus-based methods generally yielded higher recall (up to 95.2%), whereas knowledge-based methods delivered higher precision (up to 72.4%). Finally, semantic relationships accounted for roughly 20% of all identified word connections, proving that meaning-based connections add substantial value beyond exact word overlap.
These findings indicate that organizations processing short texts can achieve meaningful performance gains by integrating semantic word matching without requiring complex supervised model training. Corpus-based techniques offer the advantage of not requiring manually curated lexical databases, whereas knowledge-based techniques provide finer-grained precision. However, because semantic matching increases similarity scores across related words, it can occasionally misclassify distinct texts with overlapping themes as paraphrases.
Organizations seeking to improve text-matching pipelines should consider adopting hybrid similarity measures, balancing corpus-based methods for broad coverage and knowledge-based methods for precision. For future research and system improvements, the article recommends moving beyond the current unordered word approach to explore structured representations, such as semantic parse trees and predicate logic, to account for sentence grammar and argument roles.
The study's primary limitation is its reliance on a bag-of-words model that ignores grammatical structure and word order, along with restricting word comparisons to identical parts of speech. Confidence in the core conclusion—that semantic metrics statistically outperform standard vector baselines—is high, bounded by the dataset's human annotator agreement level of approximately 83%.
- Paper: Using Information Content to Evaluate Semantic Similarity in a Taxonomy, Philip Resnik (1995). It introduces the foundational information-content metric on taxonomy hierarchies that serves as one of the primary knowledge-based word similarity measures evaluated and combined in the paper.
- Paper: Semantic Similarity Based on Corpus Statistics and Lexical Taxonomy, Jay J. Jiang et al. (1997). It formulates the hybrid semantic similarity measure combining corpus statistics with taxonomy edge distances that the source paper directly adapts to compute pairwise word similarities.
- Paper: An Information-Theoretic Definition of Similarity, Dekang Lin (1998). It defines the core information-theoretic similarity formula used by the source paper as a principal knowledge-based and corpus-based component.
- Paper: WordNet::Similarity - Measuring the Relatedness of Concepts, Ted Pedersen et al. (2004). It details the implementation of standard WordNet-based similarity measures that the paper integrates to evaluate text snippets.
- Paper: Mining the Web for Synonyms: PMI-IR versus LSA on TOEFL, Peter D. Turney (2001). It presents Pointwise Mutual Information using search engine statistics (PMI-IR), providing the foundation for the corpus-based word similarity measures used in the source.
- Paper: Term-weighting approaches in automatic text retrieval, Gerard Salton et al. (1988). It establishes the vector space model and tf-idf weighting scheme that the source paper uses as its baseline comparison for text matching.
- Paper: Computing Semantic Relatedness Using Wikipedia-based Explicit Semantic Analysis, Evgeniy Gabrilovich et al. (2007). It advances beyond WordNet and pairwise word aggregation by introducing Explicit Semantic Analysis to represent text passages in a high-dimensional concept space.
- Paper: SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation, Daniel Cer et al. (2017). It establishes the standardized multilingual benchmark and shared evaluation framework that became the modern successor for sentence-level semantic textual similarity.
- Paper: Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Nils Reimers et al. (2019). It replaces combinatorial pairwise word similarity aggregation with dense sentence-level embeddings fine-tuned for semantic textual similarity.
- Paper: From Word Embeddings To Document Distances, Matt J. Kusner et al. (2015). It extends word-level semantic alignments between short texts into a globally optimal transportation distance metric for arbitrary documents.
- Paper: Distributed Representations of Sentences and Documents, Quoc V. Le et al. (2014). It presents an unsupervised neural approach to directly learn dense vector representations for sentences and snippets rather than aggregating lexical similarity scores.
- Paper: Skip-Thought Vectors, Ryan Kiros et al. (2015). It develops neural sentence encoders that produce fixed-dimensional sentence representations for semantic relatedness and paraphrase detection without manual feature combination.
- Paper: From Frequency to Meaning: Vector Space Models of Semantics, Peter D. Turney et al. (2010). It provides a comprehensive survey unifying vector space models of semantics and how word-level context matrices relate to broader text meaning.
