keyword
word co-occurrence statistics
Word co-occurrence statistics are quantitative measurements of how frequently pairs or groups of words appear together within a specified textual boundary, such as a sliding context window, sentence, or document. Based on the principle that words with similar meanings tend to occur in similar contexts, these counts are gathered across large text corpora to model semantic and syntactic associations between words. In natural language processing, raw co-occurrence frequencies are typically structured into word-context matrices and transformed using statistical association metrics such as pointwise mutual information or conditional probabilities. These statistics serve as the foundation for traditional count-based distributional semantic models, automated evaluation of topic model coherence, and matrix factorization methods that generate dense vector embeddings for assessing word similarity and relational analogies.
4 items

Improving Distributional Similarity with Lessons Learned from Word Embeddings
Omer Levy, Yoav Goldberg, Ido Dagan
Why you should read this
Demonstrates that the superior performance of neural word embeddings over traditional count-based models stems from hyperparameter optimizations rather than algorithmic differences, proving that applying these same tuning strategies to count-based methods eliminates the performance gap across semantic benchmarks.
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.
Added
2026-09-25

Don’t count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors
Marco Baroni, Georgiana Dinu, Germán Kruszewski
Why you should read this
Demonstrates through an extensive empirical evaluation across multiple lexical semantics benchmarks that context-predicting word embedding models consistently outperform traditional count-based distributional semantic vectors across various parameter configurations.
Context-predicting models (more commonly known as embeddings or neural language models) are the new kids on the distributional semantics block. Despite the buzz surrounding these models, the literature is still lacking a systematic comparison of the predictive models with classic, count-vector-based distributional semantic approaches. In this paper, we perform such an extensive evaluation, on a wide range of lexical semantics tasks and across many parameter settings. The results, to our own surprise, show that the buzz is fully justified, as the context-predicting models obtain a thorough and resounding victory against their count-based counterparts.
Added
2026-09-25

Optimizing Semantic Coherence in Topic Models
David Mimno, Hanna M. Wallach, Edmund Talley, Miriam Leenders, Andrew McCallum
Why you should read this
Proposes an intrinsic semantic coherence metric and a generalized Pólya urn topic model to automatically detect and eliminate low-quality, nonsensical latent topics without requiring human evaluation or external corpora.
Latent variable models have the potential to add value to large document collections by discovering interpretable, low-dimensional subspaces. In order for people to use such models, however, they must trust them. Unfortunately, typical dimensionality reduction methods for text, such as latent Dirichlet allocation, often produce low-dimensional subspaces (topics) that are obviously flawed to human domain experts. The contributions of this paper are threefold: (1) An analysis of the ways in which topics can be flawed; (2) an automated evaluation metric for identifying such topics that does not rely on human annotators or reference collections outside the training data; (3) a novel statistical topic model based on this metric that significantly improves topic quality in a large-scale document collection from the National Institutes of Health (NIH).
Added
2026-09-17

Neural Word Embedding as Implicit Matrix Factorization
Omer Levy, Yoav Goldberg
Why you should read this
Proves that popular neural word embedding models like word2vec's skip-gram with negative sampling are mathematically equivalent to factorizing a shifted pointwise mutual information matrix, bridging the theoretical gap between neural approaches and classical count-based distributional semantics.
We analyze skip-gram with negative-sampling (SGNS), a word embedding method introduced by Mikolov et al., and show that it is implicitly factorizing a word-context matrix, whose cells are the pointwise mutual information (PMI) of the respective word and context pairs, shifted by a global constant. We find that another embedding method, NCE, is implicitly factorizing a similar matrix, where each cell is the (shifted) log conditional probability of a word given its context. We show that using a sparse Shifted Positive PMI word-context matrix to represent words improves results on two word similarity tasks and one of two analogy tasks. When dense low-dimensional vectors are preferred, exact factorization with SVD can achieve solutions that are at least as good as SGNS's solutions for word similarity tasks. On analogy questions SGNS remains superior to SVD. We conjecture that this stems from the weighted nature of SGNS's factorization.
Added
2026-09-17
