Built independently by an author, for readers. Read the story and support ChapterPal

keyword

distributional similarity

Distributional similarity is a measure of the semantic resemblance between words or linguistic units based on the degree to which they appear in similar surrounding textual contexts across large bodies of text. Rooted in the linguistic principle known as the distributional hypothesis, which suggests that words with similar meanings tend to occur in similar environments, this approach quantifies semantic closeness without relying exclusively on manual dictionaries or predefined taxonomies. In computational linguistics and natural language processing, distributional similarity is typically calculated by representing vocabulary items as numerical vectors derived from word co-occurrence statistics or learned word embeddings, and then computing geometric distance or cosine similarity between those vectors in a shared mathematical space. This allows automated systems to effectively identify synonyms, assess semantic relatedness, and perform analogy detection across various text analysis tasks.

3 items

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Rada Mihalcea, Courtney Corley, Carlo Strapparava

OrganizationsDepartment of Computer ScienceFondazione Bruno KesslerIstituto per la Ricerca Scientifica e TecnologicaUniversity of North Texas

Why you should read this

Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.

This paper presents a method for measuring the semantic similarity of texts, using corpus-based and knowledge-based measures of similarity. Previous work on this problem has focused mainly on either large documents (e.g. text classification, information retrieval) or individual words (e.g. synonymy tests). Given that a large fraction of the information available today, on the Web and elsewhere, consists of short text snippets (e.g. abstracts of scientific documents, image captions, product descriptions), in this paper we focus on measuring the semantic similarity of short texts. Through experiments performed on a paraphrase data set, we show that the semantic similarity method outperforms methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.

Added

2026-09-25