keyword
word similarity tasks
Word similarity tasks are standard evaluation benchmarks in natural language processing designed to assess how accurately computational word representations capture semantic closeness between pairs of words. In these evaluations, a model calculates a mathematical similarity score, typically the cosine similarity between the vector embeddings of two words across a standardized dataset. The resulting model scores or rankings are then compared against human judgment ratings using statistical correlation metrics, such as Spearman rank correlation or Pearson correlation coefficients. By measuring how well geometric distances in a vector space reflect human linguistic intuition, these tasks serve as a foundational intrinsic method for testing and comparing traditional distributional semantic models and neural word embeddings.
3 items

Improving Distributional Similarity with Lessons Learned from Word Embeddings
Omer Levy, Yoav Goldberg, Ido Dagan
Why you should read this
Demonstrates that the superior performance of neural word embeddings over traditional count-based models stems from hyperparameter optimizations rather than algorithmic differences, proving that applying these same tuning strategies to count-based methods eliminates the performance gap across semantic benchmarks.
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.
Added
2026-09-25

Neural Word Embedding as Implicit Matrix Factorization
Omer Levy, Yoav Goldberg
Why you should read this
Proves that popular neural word embedding models like word2vec's skip-gram with negative sampling are mathematically equivalent to factorizing a shifted pointwise mutual information matrix, bridging the theoretical gap between neural approaches and classical count-based distributional semantics.
We analyze skip-gram with negative-sampling (SGNS), a word embedding method introduced by Mikolov et al., and show that it is implicitly factorizing a word-context matrix, whose cells are the pointwise mutual information (PMI) of the respective word and context pairs, shifted by a global constant. We find that another embedding method, NCE, is implicitly factorizing a similar matrix, where each cell is the (shifted) log conditional probability of a word given its context. We show that using a sparse Shifted Positive PMI word-context matrix to represent words improves results on two word similarity tasks and one of two analogy tasks. When dense low-dimensional vectors are preferred, exact factorization with SVD can achieve solutions that are at least as good as SGNS's solutions for word similarity tasks. On analogy questions SGNS remains superior to SVD. We conjecture that this stems from the weighted nature of SGNS's factorization.
Added
2026-09-17

GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, Christopher D. Manning
Why you should read this
Demonstrates how to unify the complementary strengths of global statistical methods and local context window approaches into a single model that efficiently captures meaningful semantic relationships in word vectors through weighted co-occurrence statistics rather than sparse matrix factorization or individual context windows.
Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arithmetic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word cooccurrence matrix, rather than on the entire sparse matrix or on individual context windowsinalargecorpus. Themodelproduces a vector space with meaningful substructure, as evidenced by its performance of 75% on a recent word analogy task. It also outperforms related models on similarity tasks and named entity recognition.
Added
2026-02-21
