Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Word similarity measures

Word similarity measures are computational metrics used in natural language processing to determine and quantify the degree of semantic likeness or conceptual closeness between pairs of words. These measures generally fall into two primary categories: knowledge-based approaches, which assess taxonomic distance and structural relationships within curated lexical databases or ontologies, and corpus-based or distributional approaches, which evaluate statistical co-occurrences and vector proximity across large collections of text. By determining semantic closeness rather than relying solely on exact lexical matches, word similarity measures provide an essential foundation for various automated language tasks, including information retrieval, thesaurus construction, paraphrase detection, and word sense disambiguation.

2 items

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Rada Mihalcea, Courtney Corley, Carlo Strapparava

OrganizationsDepartment of Computer ScienceFondazione Bruno KesslerIstituto per la Ricerca Scientifica e TecnologicaUniversity of North Texas

Why you should read this

Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.

This paper presents a method for measuring the semantic similarity of texts, using corpus-based and knowledge-based measures of similarity. Previous work on this problem has focused mainly on either large documents (e.g. text classification, information retrieval) or individual words (e.g. synonymy tests). Given that a large fraction of the information available today, on the Web and elsewhere, consists of short text snippets (e.g. abstracts of scientific documents, image captions, product descriptions), in this paper we focus on measuring the semantic similarity of short texts. Through experiments performed on a paraphrase data set, we show that the semantic similarity method outperforms methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.

Added

2026-09-25