word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method
Yoav GoldbergOmer Levy
Explains the mathematical derivation and intuitive rationale behind Mikolov et al.'s negative-sampling objective, providing a clear formulation of how word2vec optimizes word embeddings.
Modern natural language processing relies heavily on converting text into numerical vectors that capture word meanings. While the popular word2vec software achieved state-of-the-art results, its original mathematical explanations and design choices remained difficult to interpret. The article provides a clear, step-by-step derivation of the negative-sampling method used in word2vec and analyzes practical implementation choices embedded in the software.
The analysis demonstrates that negative sampling fundamentally reframes the learning problem. Instead of predicting a context word given a target word—which requires computationally prohibitive calculations across hundreds of thousands of vocabulary words—negative sampling converts the task into a binary classification problem. The model learns to distinguish true word-context pairs observed in the training text from randomly generated negative pairs. Without these negative samples, the model would produce a trivial solution where all vector representations collapse into identical values.
In addition to the mathematical formulation, the article highlights critical implementation details in the software that significantly alter the training data. The software utilizes a dynamic context window, choosing a random window size up to a defined maximum for each word. Furthermore, infrequent words are pruned entirely, and highly frequent words are down-sampled before generating context windows. This preprocessing step effectively expands the window size across remaining words, allowing the model to capture broader topical relationships between distant, content-heavy words.
These findings show that word2vec's high performance and computational efficiency stem not only from its classification objective, but also from subtle data preparation heuristics. The authors note, however, that a theoretical gap remains: while the system relies on the assumption that words appearing in similar contexts share similar meanings, a formal mathematical explanation for why this specific optimization objective consistently produces high-quality word representations has yet to be established.
- Paper: Distributed Representations of Words and Phrases and their Compositionality, Tomas Mikolov et al. (2013). This seminal paper introduces the negative-sampling Skip-gram objective whose mathematical derivation and theoretical rationale the source note directly sets out to explain.
- Paper: Efficient Estimation of Word Representations in Vector Space, Tomáš Mikolov et al. (2013). It introduces the Continuous Bag-of-Words and Skip-gram architectures that establish the foundational framework for word2vec representations.
- Paper: A Neural Probabilistic Language Model, Yoshua Bengio et al. (2003). This foundational work establishes continuous vector space word representations via neural language modeling upon which word2vec builds.
- Paper: From Frequency to Meaning: Vector Space Models of Semantics, Peter D. Turney et al. (2010). It provides a comprehensive foundation for vector space models and distributional semantics underlying modern word embeddings.
- Paper: Neural Word Embedding as Implicit Matrix Factorization, Omer Levy et al. (2014). It rigorously proves that the word2vec negative-sampling objective derived in the source is mathematically equivalent to implicit shifted Pointwise Mutual Information matrix factorization.
- Paper: GloVe: Global Vectors for Word Representation, Jeffrey Pennington et al. (2014). It introduces GloVe as an alternative word representation model combining global co-occurrence matrix factorization with the log-bilinear predictive strengths of word2vec.
- Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). It extends the negative-sampling skip-gram framework explained in the source by representing words as bags of character n-grams to handle subword morphology.
- Paper: Distributed Representations of Sentences and Documents, Quoc V. Le et al. (2014). It generalizes the word2vec framework from individual words to variable-length paragraphs and documents.
- Paper: DeepWalk: online learning of social representations, Bryan Perozzi et al. (2014). It adapts the Skip-gram architecture to graph structures by treating random walks as sentences to generate vertex embeddings.
- Paper: node2vec: Scalable Feature Learning for Networks, Aditya Grover et al. (2016). It builds directly on word2vec's Skip-gram with negative sampling to learn versatile node representations in networks using biased random walks.
- Paper: LINE: Large-scale Information Network Embedding, Jian Tang et al. (2015). It applies the negative-sampling optimization objective from word2vec to learn scalable node embeddings that preserve first- and second-order network proximity.
- Paper: Convolutional Neural Networks for Sentence Classification, Yoon Kim (2014). It demonstrates how pre-trained word2vec embeddings can be directly leveraged as feature inputs to boost sentence classification performance in convolutional neural networks.
- Paper: From Word Embeddings To Document Distances, Matt J. Kusner et al. (2015). It utilizes pre-trained word2vec embeddings to formulate an Earth Mover's Distance-based metric for computing distances between entire text documents.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). It investigates and develops debiasing techniques for societal stereotypes encoded within geometry of pre-trained word2vec vector spaces.
