keyword
distributional similarity
Distributional similarity is a measure of the semantic resemblance between words or linguistic units based on the degree to which they appear in similar surrounding textual contexts across large bodies of text. Rooted in the linguistic principle known as the distributional hypothesis, which suggests that words with similar meanings tend to occur in similar environments, this approach quantifies semantic closeness without relying exclusively on manual dictionaries or predefined taxonomies. In computational linguistics and natural language processing, distributional similarity is typically calculated by representing vocabulary items as numerical vectors derived from word co-occurrence statistics or learned word embeddings, and then computing geometric distance or cosine similarity between those vectors in a shared mathematical space. This allows automated systems to effectively identify synonyms, assess semantic relatedness, and perform analogy detection across various text analysis tasks.
3 items

Improving Distributional Similarity with Lessons Learned from Word Embeddings
Omer Levy, Yoav Goldberg, Ido Dagan
Why you should read this
Demonstrates that the superior performance of neural word embeddings over traditional count-based models stems from hyperparameter optimizations rather than algorithmic differences, proving that applying these same tuning strategies to count-based methods eliminates the performance gap across semantic benchmarks.
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.
Added
2026-09-25

Corpus-based and Knowledge-based Measures of Text Semantic Similarity
Rada Mihalcea, Courtney Corley, Carlo Strapparava
Why you should read this
Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.
This paper presents a method for measuring the semantic similarity of texts, using corpus-based and knowledge-based measures of similarity. Previous work on this problem has focused mainly on either large documents (e.g. text classification, information retrieval) or individual words (e.g. synonymy tests). Given that a large fraction of the information available today, on the Web and elsewhere, consists of short text snippets (e.g. abstracts of scientific documents, image captions, product descriptions), in this paper we focus on measuring the semantic similarity of short texts. Through experiments performed on a paraphrase data set, we show that the semantic similarity method outperforms methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.
Added
2026-09-25

Computing Semantic Relatedness Using Wikipedia-based Explicit Semantic Analysis
Evgeniy Gabrilovich, Shaul Markovitch
Why you should read this
Shows how Explicit Semantic Analysis turns Wikipedia concepts into interpretable text vectors and substantially improves agreement with human semantic-relatedness judgments.
Computing semantic relatedness of natural language texts requires access to vast amounts of common-sense and domain-specific world knowledge. We propose Explicit Semantic Analysis (ESA), a novel method that represents the meaning of texts in a high-dimensional space of concepts derived from Wikipedia. We use machine learning techniques to explicitly represent the meaning of any text as a weighted vector of Wikipedia-based concepts. Assessing the relatedness of texts in this space amounts to comparing the corresponding vectors using conventional metrics (e.g., cosine). Compared with the previous state of the art, using ESA results in substantial improvements in correlation of computed relatedness scores with human judgments: from r = 0.56 to 0.75 for individual words and from r = 0.60 to 0.72 for texts. Importantly, due to the use of natural concepts, the ESA model is easy to explain to human users.
Added
2026-09-14
