Built independently by an author, for readers. Read the story and support ChapterPal

keyword

word co-occurrence statistics

Word co-occurrence statistics are quantitative measurements of how frequently pairs or groups of words appear together within a specified textual boundary, such as a sliding context window, sentence, or document. Based on the principle that words with similar meanings tend to occur in similar contexts, these counts are gathered across large text corpora to model semantic and syntactic associations between words. In natural language processing, raw co-occurrence frequencies are typically structured into word-context matrices and transformed using statistical association metrics such as pointwise mutual information or conditional probabilities. These statistics serve as the foundation for traditional count-based distributional semantic models, automated evaluation of topic model coherence, and matrix factorization methods that generate dense vector embeddings for assessing word similarity and relational analogies.

4 items

Neural Word Embedding as Implicit Matrix Factorization

Neural Word Embedding as Implicit Matrix Factorization

Omer Levy, Yoav Goldberg

OrganizationsBar-Ilan University

Why you should read this

Proves that popular neural word embedding models like word2vec's skip-gram with negative sampling are mathematically equivalent to factorizing a shifted pointwise mutual information matrix, bridging the theoretical gap between neural approaches and classical count-based distributional semantics.

We analyze skip-gram with negative-sampling (SGNS), a word embedding method introduced by Mikolov et al., and show that it is implicitly factorizing a word-context matrix, whose cells are the pointwise mutual information (PMI) of the respective word and context pairs, shifted by a global constant. We find that another embedding method, NCE, is implicitly factorizing a similar matrix, where each cell is the (shifted) log conditional probability of a word given its context. We show that using a sparse Shifted Positive PMI word-context matrix to represent words improves results on two word similarity tasks and one of two analogy tasks. When dense low-dimensional vectors are preferred, exact factorization with SVD can achieve solutions that are at least as good as SGNS's solutions for word similarity tasks. On analogy questions SGNS remains superior to SVD. We conjecture that this stems from the weighted nature of SGNS's factorization.

Added

2026-09-17