Built independently by an author, for readers. Read the story and support ChapterPal

keyword

word similarity tasks

Word similarity tasks are standard evaluation benchmarks in natural language processing designed to assess how accurately computational word representations capture semantic closeness between pairs of words. In these evaluations, a model calculates a mathematical similarity score, typically the cosine similarity between the vector embeddings of two words across a standardized dataset. The resulting model scores or rankings are then compared against human judgment ratings using statistical correlation metrics, such as Spearman rank correlation or Pearson correlation coefficients. By measuring how well geometric distances in a vector space reflect human linguistic intuition, these tasks serve as a foundational intrinsic method for testing and comparing traditional distributional semantic models and neural word embeddings.

3 items

Neural Word Embedding as Implicit Matrix Factorization

Neural Word Embedding as Implicit Matrix Factorization

Omer Levy, Yoav Goldberg

OrganizationsBar-Ilan University

Why you should read this

Proves that popular neural word embedding models like word2vec's skip-gram with negative sampling are mathematically equivalent to factorizing a shifted pointwise mutual information matrix, bridging the theoretical gap between neural approaches and classical count-based distributional semantics.

We analyze skip-gram with negative-sampling (SGNS), a word embedding method introduced by Mikolov et al., and show that it is implicitly factorizing a word-context matrix, whose cells are the pointwise mutual information (PMI) of the respective word and context pairs, shifted by a global constant. We find that another embedding method, NCE, is implicitly factorizing a similar matrix, where each cell is the (shifted) log conditional probability of a word given its context. We show that using a sparse Shifted Positive PMI word-context matrix to represent words improves results on two word similarity tasks and one of two analogy tasks. When dense low-dimensional vectors are preferred, exact factorization with SVD can achieve solutions that are at least as good as SGNS's solutions for word similarity tasks. On analogy questions SGNS remains superior to SVD. We conjecture that this stems from the weighted nature of SGNS's factorization.

Added

2026-09-17

GloVe: Global Vectors for Word Representation

GloVe: Global Vectors for Word Representation

Jeffrey Pennington, Richard Socher, Christopher D. Manning

OrganizationsStanford University

Why you should read this

Demonstrates how to unify the complementary strengths of global statistical methods and local context window approaches into a single model that efficiently captures meaningful semantic relationships in word vectors through weighted co-occurrence statistics rather than sparse matrix factorization or individual context windows.

Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arithmetic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word cooccurrence matrix, rather than on the entire sparse matrix or on individual context windowsinalargecorpus. Themodelproduces a vector space with meaningful substructure, as evidenced by its performance of 75% on a recent word analogy task. It also outperforms related models on similarity tasks and named entity recognition.

Added

2026-02-21