Improving Distributional Similarity with Lessons Learned from Word Embeddings

Omer LevyYoav GoldbergIdo Dagan

article2015TACL1,409 citations

Demonstrates that the superior performance of neural word embeddings over traditional count-based models stems from hyperparameter optimizations rather than algorithmic differences, proving that applying these same tuning strategies to count-based methods eliminates the performance gap across semantic benchmarks.

Listen

Recent advancements in natural language processing have popularized neural-network-inspired word embedding algorithms, with widespread claims that these modern prediction-based techniques inherently outperform traditional statistical count-based models. Determining whether this superior performance stems from core architectural breakthroughs or other engineering factors is crucial for teams allocating computational budgets and selecting text processing frameworks. The article evaluates four leading word representation methods across identical settings to demonstrate how engineering design choices and system configuration settings—collectively termed hyperparameters—drive observed performance gains across models.

To perform this evaluation, the article compared two traditional count-based models (positive pointwise mutual information matrices and singular value decomposition) against two neural embedding models (skip-gram with negative sampling and Global Vectors). The analysis examined 672 configurations across eight standard word similarity and analogy benchmark datasets using an English Wikipedia corpus of 1.5 billion tokens, with additional scale validation on a 10.5-billion-word corpus. The investigation isolated pre-processing, association metric, and post-processing adjustments—such as context distribution smoothing and dynamic context windows—and transferred them across all applicable approaches.

Key findings show that hyperparameter tuning accounts for the majority of reported performance advantages rather than the underlying embedding algorithms themselves. Once configuration choices are controlled and adapted across methods, the article found no global advantage for any single approach; traditional count-based methods perform comparably to neural embeddings on word similarity tasks, with singular value decomposition frequently matching or exceeding skip-gram performance. Adapting context distribution smoothing to pointwise mutual information measures consistently increased accuracy by over 3 points per task on average by mitigating bias toward rare words. Furthermore, proper hyperparameter optimization frequently produced larger performance gains—up to 15.7 percentage points over default baselines—than switching algorithmic architectures or significantly increasing training corpus size.

These findings indicate that organizations do not necessarily need to migrate to complex neural pipelines to achieve competitive language understanding accuracy. Teams can avoid costly model replacements by transferring lightweight optimization techniques into existing traditional infrastructures. However, skip-gram with negative sampling remains the most computationally efficient and memory-friendly baseline across diverse tasks, training in roughly half a day on massive text corpora where competing dense factorizations like GloVe require several days or exceed system memory.

For immediate implementation, organizations should apply context distribution smoothing when generating association measures and utilize symmetric singular value decomposition configurations rather than standard textbook factorizations, which severely degrade accuracy. Practitioners using skip-gram models should tune toward larger numbers of negative samples and evaluate additive word-plus-context vector representations. While these conclusions hold across extensive linguistic benchmarks and English corpora, engineering teams should evaluate hyperparameter tuning using separate cross-validation splits for their specific downstream applications before committing large computational resources.

Levy et al (2015).pdf
Cover for Improving Distributional Similarity with Lessons Learned from Word Embeddings

Abstract

Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Explicit Representations (PPMI Matrix)
  • 2.2 Singular Value Decomposition (SVD)
  • 2.3 Skip-Grams with Negative Sampling (SGNS)
  • 2.4 Global Vectors (GloVe)
  • 3 Transferable Hyperparameters
  • 3.1 Pre-processing Hyperparameters
  • 3.2 Association Metric Hyperparameters
  • 3.3 Post-processing Hyperparameters
  • 4 Experimental Setup
  • 4.1 Hyperparameter Space
  • 4.2 Word Representations
  • 4.3 Test Datasets
  • 5 Results
  • 5.1 Hyperparameters vs Algorithms
  • 5.2 Hyperparameters vs Big Data
  • 5.3 Re-evaluating Prior Claims
  • 5.4 Comparison with CBOW
  • 6 Hyperparameter Analysis
  • 6.1 Harmful Configurations
  • 6.2 Beneficial Configurations
  • 7 Practical Recommendations
  • 8 Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Context Distribution Smoothing for Pointwise Mutual Information

    model/method

    Pointwise Mutual Information (PMI) exhibits an inherent bias toward infrequent context words because the estimated marginal context probability in the denominator is extremely small for rare contexts. Context Distribution Smoothing (CDS) transfers the smoothing applied to negative sampling distributions in skip-gram neural embeddings (SGNS) to explicit PMI calculations by raising context frequencies to an exponent α∈(0,1]\alpha \in (0, 1] (typically α=0.75\alpha = 0.75).

    Given a corpus of observed word-context pairs DD, where #(w,c)\#(w, c) denotes the co-occurrence count of target word ww and context word cc, #(w)=∑c′∈VC#(w,c′)\#(w) = \sum_{c' \in V_C} \#(w, c') is the marginal word count, and #(c)=∑w′∈VW#(w′,c)\#(c) = \sum_{w' \in V_W} \#(w', c) is the marginal context count, smoothed Pointwise Mutual Information is defined as: PMIα(w,c)=log⁡P^(w,c)P^(w)P^α(c)\text{PMI}_\alpha(w, c) = \log \frac{\hat{P}(w, c)}{\hat{P}(w)\hat{P}_\alpha(c)} where P^(w,c)=#(w,c)∣D∣\hat{P}(w, c) = \frac{\#(w, c)}{|D|}, P^(w)=#(w)∣D∣\hat{P}(w) = \frac{\#(w)}{|D|}, and the smoothed context unigram distribution is: P^α(c)=#(c)α∑c′∈VC#(c′)α\hat{P}_\alpha(c) = \frac{\#(c)^\alpha}{\sum_{c' \in V_C} \#(c')^\alpha} When α=1\alpha = 1, the expression yields standard unsmoothed PMI. When α=0.75\alpha = 0.75, the probability of sampling rare contexts increases (since P^α(c)>P^(c)\hat{P}_\alpha(c) > \hat{P}(c) for infrequent cc), which lowers the PMI value for word-context pairs involving rare contexts. Applying non-negative clipping yields smoothed Positive Pointwise Mutual Information (smoothed PPMI): PPMIα(w,c)=max⁡(PMIα(w,c),0)\text{PPMI}_\alpha(w, c) = \max(\text{PMI}_\alpha(w, c), 0) This modification consistently improves semantic similarity and analogy recovery across explicit and low-rank factorized representations.

  2. Knowl 2 — Shifted Positive Pointwise Mutual Information

    model/method

    Skip-gram with negative sampling (SGNS) implicitly factorizes a word-context matrix whose entries equal Pointwise Mutual Information (PMI) shifted by −log⁡k-\log k, where kk is the number of negative samples drawn per positive observation. This shift parameter can be transferred directly to count-based distributional models as Shifted Positive Pointwise Mutual Information (SPPMI).

    Given target word w∈VWw \in V_W and context word c∈VCc \in V_C, with empirical joint count #(w,c)\#(w, c), marginal word count #(w)\#(w), marginal context count #(c)\#(c), and total pair count ∣D∣|D|, standard PMI is: PMI(w,c)=log⁡#(w,c)⋅∣D∣#(w)⋅#(c)\text{PMI}(w, c) = \log \frac{\#(w, c) \cdot |D|}{\#(w) \cdot \#(c)} For an integer shift parameter k≥1k \ge 1, Shifted PPMI is defined as: SPPMIk(w,c)=max⁡(PMI(w,c)−log⁡k,0)\text{SPPMI}_k(w, c) = \max(\text{PMI}(w, c) - \log k, 0) The hyperparameter kk acts as a prior threshold: only word-context pairs with mutual information exceeding log⁡k\log k retain positive values, while weaker associations are set to 0. Setting k>1k > 1 creates sparser explicit representations with higher-confidence associations, which improves performance on word similarity benchmarks for explicit PPMI and SGNS.

  3. Knowl 3 — Symmetric Eigenvalue Weighting in Truncated SVD for Word Embeddings

    model/method

    Truncated Singular Value Decomposition (SVD) factors a high-dimensional word-context association matrix MM (such as a PPMI matrix) of rank rr into Md=UdΣdVd⊤M_d = U_d \Sigma_d V_d^\top, where Ud∈R∣VW∣×dU_d \in \mathbb{R}^{|V_W| \times d} and Vd∈R∣VC∣×dV_d \in \mathbb{R}^{|V_C| \times d} have orthonormal columns, Σd∈Rd×d\Sigma_d \in \mathbb{R}^{d \times d} is a diagonal matrix containing the dd largest singular values in descending order, and d<rd < r.

    The standard textbook formulation in distributional semantics (e.g., Latent Semantic Analysis) assigns the full eigenvalue matrix to the word embeddings: W=UdΣd,C=VdW = U_d \Sigma_d, \quad C = V_d This construction produces an asymmetric representation where context vectors in CC are orthonormal but word vectors in WW are not. Parameterizing the factorization by an eigenvalue exponent p∈[0,1]p \in [0, 1] yields the generalized formulation: W=UdΣdp,C=VdΣd1−pW = U_d \Sigma_d^p, \quad C = V_d \Sigma_d^{1-p} Two symmetric alternatives are:

    1. Split eigenvalue weighting (p=0.5p = 0.5): W=UdΣd,C=VdΣdW = U_d \sqrt{\Sigma_d}, \quad C = V_d \sqrt{\Sigma_d}
    2. Eigenvalue discarding (p=0p = 0): W=Ud,C=VdW = U_d, \quad C = V_d In word similarity evaluations, the standard setting (p=1p = 1) performs substantially worse than symmetric variants (p=0.5p = 0.5 or p=0p = 0), frequently underperforming by over 15 absolute points in Spearman's ρ\rho correlation.
  4. Knowl 4 — First- and Second-Order Similarity Decomposition via Word and Context Vector Addition

    theoretical result

    In distributional models where words and contexts share a unified vocabulary (VW=VCV_W = V_C) and yield distinct word embeddings w⃗∈Rd\vec{w} \in \mathbb{R}^d and context embeddings c⃗∈Rd\vec{c} \in \mathbb{R}^d (such as SGNS, SVD, and GloVe), a word xx can be represented post-training by adding its word and context vectors: v⃗x=w⃗x+c⃗x\vec{v}_x = \vec{w}_x + \vec{c}_x Assuming word and context vectors are normalized to unit length after training (∥w⃗x∥=∥c⃗x∥=1\|\vec{w}_x\| = \|\vec{c}_x\| = 1), the cosine similarity between two word representations v⃗x\vec{v}_x and v⃗y\vec{v}_y expands to: cos⁡(v⃗x,v⃗y)=w⃗x⋅w⃗y+c⃗x⋅c⃗y+w⃗x⋅c⃗y+c⃗x⋅w⃗y2w⃗x⋅c⃗x+1w⃗y⋅c⃗y+1\cos(\vec{v}_x, \vec{v}_y) = \frac{\vec{w}_x \cdot \vec{w}_y + \vec{c}_x \cdot \vec{c}_y + \vec{w}_x \cdot \vec{c}_y + \vec{c}_x \cdot \vec{w}_y}{2 \sqrt{\vec{w}_x \cdot \vec{c}_x + 1} \sqrt{\vec{w}_y \cdot \vec{c}_y + 1}} This formulation combines two distinct semantic notions:

    1. Second-order similarity (w⃗x⋅w⃗y+c⃗x⋅c⃗y\vec{w}_x \cdot \vec{w}_y + \vec{c}_x \cdot \vec{c}_y): measures whether xx and yy appear in similar context distributions, reflecting Harris's distributional hypothesis.
    2. First-order similarity (w⃗x⋅c⃗y+c⃗x⋅w⃗y\vec{w}_x \cdot \vec{c}_y + \vec{c}_x \cdot \vec{w}_y): measures whether xx and yy directly co-occur as target and context.

    The denominator normalizes by the reflective first-order self-similarities w⃗x⋅c⃗x\vec{w}_x \cdot \vec{c}_x and w⃗y⋅c⃗y\vec{w}_y \cdot \vec{c}_y. Applying the additive w⃗+c⃗\vec{w} + \vec{c} representation is computationally free (requiring no model retraining) and frequently boosts semantic similarity scores for SGNS and GloVe, although it can degrade performance on morpho-syntactic analogy tasks.

  5. Knowl 5 — Dynamic Context Windows and Subsampling in Matrix-Based Distributional Models

    model/method

    System design hyperparameters introduced in neural word embedding implementations transfer directly to count-based and factorization-based matrix representations:

    1. Dynamic Context Windows: Traditional count models use an unweighted window of fixed length LL. Dynamic context windowing weights context words inversely to their distance from the focus word. In skip-gram implementations, this is realized by uniformly sampling the window radius l∼U{1,L}l \sim \mathcal{U}\{1, L\} for each token occurrence. For an explicit co-occurrence matrix, this stochastic sampling is equivalent in expectation to a deterministic linear decay where a context token at distance k≤Lk \le L from the target receives weight (L−k+1)/L(L - k + 1) / L.

    2. Frequent Word Subsampling: High-frequency words (such as function words) are diluted by discarding token occurrences prior to context pair extraction with probability: p=1−tfp = 1 - \sqrt{\frac{t}{f}} where f=#(w)/∣D∣f = \#(w) / |D| is the word's relative frequency in the corpus and tt is a frequency threshold (e.g., t=10−5t = 10^{-5}). In 'dirty' subsampling, tokens are eliminated from the running text stream before defining context windows of size LL, effectively widening the context reach across discarded tokens.

    Both methods reduce computational requirements (producing fewer non-zero cells in co-occurrence matrices and fewer SGD updates in prediction models) and improve semantic vector quality across count-based and neural representations.

  6. Knowl 6 — Performance Parity Between Count-Based and Prediction-Based Word Representations

    empirical result

    When hyperparameters (context distribution smoothing, dynamic window weighting, subsampling, shift parameter, and eigenvalue scaling) are held constant and tuned under 2-fold cross-validation across 6 word similarity benchmarks (WordSim-353 Similarity, WordSim-353 Relatedness, MEN, Mechanical Turk, Rare Words, SimLex-999) and 2 analogy benchmarks (Google Analogy, MSR Analogy), neural prediction-based embeddings (SGNS, GloVe) do not demonstrate an inherent global advantage over traditional count-based representations (PPMI, SVD).

    Key empirical findings:

    1. On word similarity tasks, SVD achieves average performance comparable to or higher than SGNS at context window sizes 2 and 5 (e.g., on SimLex-999 with window size 2, SVD achieves Spearman's ρ=0.425\rho = 0.425 and SGNS achieves 0.4330.433; on Rare Words, SVD achieves 0.5080.508 while SGNS achieves 0.4490.449).
    2. On semantic analogy recovery within the Google dataset, explicit PPMI achieves performance close to SGNS (0.677 vs. 0.689 with 3CosMul at window size 2).
    3. The only benchmark where SGNS substantially outperforms count-based methods is the MSR morpho-syntactic analogy dataset (SGNS achieves 0.644 vs. PPMI's 0.535 and SVD's 0.468 at window size 2).
    4. Hyperparameter tuning produces performance swings of up to 15.7 points within a single algorithm, exceeding the typical performance gap between different algorithmic methods.
  7. Knowl 7 — Cross-Validated Performance of Distributional Representations Across Similarity and Analogy Benchmarks

    data/table

    The table below presents the 2-fold cross-validated performance of Positive Pointwise Mutual Information (PPMI), Singular Value Decomposition (SVD), Skip-Gram with Negative Sampling (SGNS), and Global Vectors (GloVe) across 6 word similarity datasets (Spearman's ρ\rho) and 2 analogy datasets (accuracy using 3CosAdd / 3CosMul) for context window sizes win∈{2,5,10}win \in \{2, 5, 10\} on a 1.5-billion-token Wikipedia corpus. Large-scale (LS) models trained on a 10.5-billion-token corpus are included for comparison.

    win Method WS-Sim WS-Rel MEN M. Turk Rare Words SimLex Google Add / Mul MSR Add / Mul
    2 PPMI .732 .699 .744 .654 .457 .382 .552 / .677 .306 / .535
    2 SVD .772 .671 .777 .647 .508 .425 .554 / .591 .408 / .468
    2 SGNS .789 .675 .773 .661 .449 .433 .676 / .689 .617 / .644
    2 GloVe .720 .605 .728 .606 .389 .388 .649 / .666 .540 / .591
    5 PPMI .732 .706 .738 .668 .442 .360 .518 / .649 .277 / .467
    5 SVD .764 .679 .776 .639 .499 .416 .532 / .569 .369 / .424
    5 SGNS .772 .690 .772 .663 .454 .403 .692 / .714 .605 / .645
    5 GloVe .745 .617 .746 .631 .416 .389 .700 / .712 .541 / .599
    10 PPMI .735 .701 .741 .663 .235 .336 .532 / .605 .249 / .353
    10 SVD .766 .681 .770 .628 .312 .419 .526 / .562 .356 / .406
    10 SGNS .794 .700 .775 .678 .281 .422 .694 / .710 .520 / .557
    10 GloVe .746 .643 .754 .616 .266 .375 .702 / .712 .463 / .519
    10 SGNS-LS .766 .681 .781 .689 .451 .414 .739 / .758 .690 / .729
    10 GloVe-LS .678 .624 .752 .639 .361 .371 .732 / .750 .628 / .685

    The cross-validated results demonstrate that no single method consistently outperforms all others across benchmarks. SVD performs best on Rare Words (.508 at win=2) and WordSim Relatedness (.679 at win=5), whereas SGNS excels on syntactic analogies in MSR (.644 with Mul at win=2). The realistic cross-validation scores fall within ~1% of oracle upper bounds.

  8. Knowl 8 — Detrimental Impact of Shifted PPMI on Truncated SVD

    empirical result

    Applying shifted Pointwise Mutual Information with shift parameter k>1k > 1, defined as SPPMI(w,c)=max⁡(PMI(w,c)−log⁡k,0)\text{SPPMI}(w, c) = \max(\text{PMI}(w, c) - \log k, 0), consistently degrades the performance of Truncated Singular Value Decomposition (SVD) across all semantic similarity and analogy tasks.

    Evaluating the difference between the best achievable SVD configurations with k>1k > 1 versus k=1k = 1 reveals substantial performance penalties:

    • WordSim Similarity: −1.7%-1.7\%
    • WordSim Relatedness: −2.2%-2.2\%
    • MEN: −1.9%-1.9\%
    • Mechanical Turk: −4.6%-4.6\%
    • Rare Words: −3.4%-3.4\%
    • SimLex-999: −3.5%-3.5\%
    • Google Analogy (3CosMul): −13.9%-13.9\%
    • MSR Analogy (3CosMul): −14.9%-14.9\%

    This failure occurs because SVD optimizes an unweighted L2L_2 reconstruction loss over all matrix entries without distinguishing observed from unobserved pairs. Shifting by log⁡k\log k converts low-positive PMI values to zero; as the proportion of zero-valued cells increases, SVD's unweighted objective forces the rank-dd factorization toward the trivial all-zero matrix. In contrast, SGNS employs a weighted, sigmoidal loss where unobserved pairs incur negligible cost, allowing it to benefit from k>1k > 1.

  9. Knowl 9 — Controlled Comparison of GloVe, SGNS, and Analogy Solvers

    empirical result

    When evaluated under controlled conditions with identical training data and hyperparameter search spaces:

    1. SGNS vs. GloVe: SGNS outperforms GloVe across every evaluated task in 2-fold cross-validation. At context window size 2, SGNS achieves higher scores on WordSim Similarity (0.789 vs. 0.720), WordSim Relatedness (0.675 vs. 0.605), MEN (0.773 vs. 0.728), SimLex-999 (0.433 vs. 0.388), Google Analogy with 3CosMul (0.689 vs. 0.666), and MSR Analogy with 3CosMul (0.644 vs. 0.591). On a 10.5-billion-word corpus, SGNS-LS similarly outperforms GloVe-LS across all tasks (e.g., 0.758 vs. 0.750 on Google Mul and 0.729 vs. 0.685 on MSR Mul).

    2. 3CosMul vs. 3CosAdd: The multiplicative analogy solver 3CosMul, arg⁡max⁡b∗∈VW∖{a∗,b,a}cos⁡(b∗,a∗)⋅cos⁡(b∗,b)cos⁡(b∗,a)+ε\arg\max_{b^* \in V_W \setminus \{a^*, b, a\}} \frac{\cos(b^*, a^*) \cdot \cos(b^*, b)}{\cos(b^*, a) + \varepsilon} consistently and strictly outperforms the additive solver 3CosAdd, arg⁡max⁡b∗∈VW∖{a∗,b,a}(cos⁡(b∗,a∗)−cos⁡(b∗,a)+cos⁡(b∗,b))\arg\max_{b^* \in V_W \setminus \{a^*, b, a\}} (\cos(b^*, a^*) - \cos(b^*, a) + \cos(b^*, b)) across all four representation methods (PPMI, SVD, SGNS, GloVe) on both the Google and MSR analogy datasets (e.g., lifting PPMI Google accuracy from 0.552 to 0.677 at window size 2, and SVD MSR accuracy from 0.408 to 0.468).

  10. Knowl 10 — Practical Rules of Thumb for Distributional Vector Space Modeling

    model/method

    Empirical analysis across eight semantic similarity and analogy benchmarks yields five actionable guidelines for training distributional vector spaces:

    1. Context Distribution Smoothing: Always apply context distribution smoothing (CDS with α=0.75\alpha = 0.75) when computing PPMI or training SGNS. It consistently boosts performance across tasks with minimal risk (e.g., improving PPMI by +9.2% on MSR analogies, +3.5% on Rare Words, and +2.8% on WordSim Relatedness).
    2. SVD Eigenvalue Weighting: Never use the textbook SVD formulation (p=1p = 1, W=UdΣdW = U_d \Sigma_d). Instead, use a symmetric decomposition (p=0.5p = 0.5, W=UdΣdW = U_d \sqrt{\Sigma_d}) or eigenvalue-free projection (p=0p = 0, W=UdW = U_d).
    3. Shifted PMI (kk): Use multiple negative samples (k=5k = 5 or k=15k = 15) for SGNS. For explicit PPMI, k>1k > 1 improves word similarity but harms analogy detection. For SVD, always retain k=1k = 1.
    4. Context Vector Addition (w⃗+c⃗\vec{w} + \vec{c}): For SGNS and GloVe, evaluate adding context vectors post-training as a computationally cost-free step that provides substantial improvements on semantic similarity datasets, while recognizing it can reduce syntactic analogy accuracy.
    5. Baseline Selection: SGNS serves as a robust, computationally efficient baseline: it is the fastest to train, consumes the least disk space and memory, and avoids catastrophic failure under varied hyperparameter configurations.

Coverage note — Deliberately omitted CBOW architectural details and the clean versus dirty subsampling comparison as secondary implementation nuances that do not alter the main conclusions on hyperparameter transferability.

References

  1. 1.Eneko Agirre, Enrique Alfonseca, Keith Hall, Jana Kravalova, Marius Pasca, and Aitor Soroa. 2009. A study on similarity and relatedness using distributional and wordnet-based approaches. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 19–27, Boulder, Colorado, June. Association for Computational Linguistics.
  2. 2.Marco Baroni and Alessandro Lenci. 2010. Distributional memory: A general framework for corpus-based semantics. Computational Linguistics, 36(4):673–721.
  3. 3.Marco Baroni, Georgiana Dinu, and Germán Kruszewski. 2014. Dont count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 238–247, Baltimore, Maryland, June. Association for Computational Linguistics.
  4. 4.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155.
  5. 5.Elia Bruni, Gemma Boleda, Marco Baroni, and Nam Khanh Tran. 2012. Distributional semantics in technicolor. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 136–145, Jeju Island, Korea, July. Association for Computational Linguistics.
  6. 6.John A Bullinaria and Joseph P Levy. 2007. Extracting semantic representations from word co-occurrence statistics: a computational study. Behavior Research Methods, 39(3):510–526.
  7. 7.John A Bullinaria and Joseph P Levy. 2012. Extracting semantic representations from word co-occurrence statistics: Stop-lists, stemming, and SVD. Behavior Research Methods, 44(3):890–907.
  8. 8.John Caron. 2001. Experiments with LSA scoring: optimal rank and basis. In Proceedings of the SIAM Computational Information Retrieval Workshop, pages 157–169.
  9. 9.Kenneth Ward Church and Patrick Hanks. 1990. Word association norms, mutual information, and lexicography. Computational Linguistics, 16(1):22–29.
  10. 10.Ronan Collobert and Jason Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of the 25th International Conference on Machine Learning, pages 160–167.
  11. 11.Scott C. Deerwester, Susan T. Dumais, Thomas K. Landauer, George W. Furnas, and Richard A. Harshman. 1990. Indexing by latent semantic analysis. JASIS, 41(6):391–407.
  12. 12.C Eckart and G Young. 1936. The approximation of one matrix by another of lower rank. Psychometrika, 1:211–218.
  13. 13.Roi Reichart Felix Hill and Anna Korhonen. 2014. Simlex-999: Evaluating semantic models with (genuine) similarity estimation. arXiv preprint arXiv:1408.3456.
  14. 14.Adriano Ferraresi, Eros Zanchetta, Marco Baroni, and Silvia Bernardini. 2008. Introducing and evaluating ukwac, a very large web-derived corpus of English. In Proceedings of the 4th Web as Corpus Workshop (WAC-4), pages 47–54.
  15. 15.Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2002. Placing search in context: The concept revisited. ACM Transactions on Information Systems, 20(1):116–131.
  16. 16.Yoav Goldberg and Omer Levy. 2014. word2vec explained: deriving Mikolov et al.’s negative-sampling word-embedding method. arXiv preprint arXiv:1402.3722.
  17. 17.Zellig Harris. 1954. Distributional structure. Word, 10(23):146–162.
  18. 18.Omer Levy and Yoav Goldberg. 2014a. Dependency-based word embeddings. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 302–308, Baltimore, Maryland.
  19. 19.Omer Levy and Yoav Goldberg. 2014b. Linguistic regularities in sparse and explicit word representations. In Proceedings of the Eighteenth Conference on Computational Natural Language Learning, pages 171–180, Baltimore, Maryland.
  20. 20.Omer Levy and Yoav Goldberg. 2014c. Neural word embeddings as implicit matrix factorization. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pages 2177–2185.
  21. 21.Minh-Thang Luong, Richard Socher, and Christopher D. Manning. 2013. Better word representations with recursive neural networks for morphology. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pages 104–113, Sofia, Bulgaria, August. Association for Computational Linguistics.
  22. 22.Oren Melamud, Ido Dagan, Jacob Goldberger, Idan Szpektor, and Deniz Yuret. 2014. Probabilistic modeling of joint-context in distributional similarity. In Proceedings of the Eighteenth Conference on Computational Natural Language Learning, pages 181–190, Baltimore, Maryland, June. Association for Computational Linguistics.
  23. 23.Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. In Proceedings of the International Conference on Learning Representations (ICLR).
  24. 24.Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems, pages 3111–3119.
  25. 25.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013c. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 746–751.
  26. 26.Sebastian Padó and Mirella Lapata. 2007. Dependency-based construction of semantic space models. Computational Linguistics, 33(2):161–199.
  27. 27.Patrick Pantel and Dekang Lin. 2002. Discovering word senses from text. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 613–619. ACM.
  28. 28.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar, October. Association for Computational Linguistics.
  29. 29.Kira Radinsky, Eugene Agichtein, Evgeniy Gabrilovich, and Shaul Markovitch. 2011. A word at a time: Computing word relatedness using temporal semantic analysis. In Proceedings of the 20th international conference on World wide web, pages 337–346. ACM.
  30. 30.Magnus Sahlgren. 2006. The Word-Space Model. Ph.D. thesis, Stockholm University.
  31. 31.Peter D. Turney and Michael L. Littman. 2003. Measuring praise and criticism: Inference of semantic orientation from association. Transactions on Information Systems, 21(4):315–346.
  32. 32.Peter D. Turney and Patrick Pantel. 2010. From frequency to meaning: Vector space models of semantics. Journal of Artificial Intelligence Research, 37(1):141–188.
  33. 33.Peter D. Turney. 2012. Domain and function: A dual-space model of semantic relations and compositions. Journal of Artificial Intelligence Research, 44:533–585.
  34. 34.Torsten Zesch, Christof Müller, and Iryna Gurevych. 2008. Using wiktionary for computing semantic relatedness. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 2, AAAI’08, pages 861–866. AAAI Press.

Citation

MLA
Levy, O., et al. “Improving Distributional Similarity with Lessons Learned from Word Embeddings”. Transactions of the Association for Computational Linguistics, 2015, pp. 211–25, https://doi.org/10.1162/tacl_a_00134.
APA
Levy, O., Goldberg, Y., & Dagan, I. (2015). Improving Distributional Similarity with Lessons Learned from Word Embeddings. Transactions of the Association for Computational Linguistics, 211–225. https://doi.org/10.1162/tacl_a_00134
Chicago
Levy, O., Y. Goldberg, and I. Dagan. 2015. “Improving Distributional Similarity with Lessons Learned from Word Embeddings”. Transactions of the Association for Computational Linguistics, 211–25. https://doi.org/10.1162/tacl_a_00134.
Harvard
Levy, O., Goldberg, Y. and Dagan, I. (2015) “Improving Distributional Similarity with Lessons Learned from Word Embeddings”, Transactions of the Association for Computational Linguistics. Association for Computational Linguistics, pp. 211–225. Available at: https://doi.org/10.1162/tacl_a_00134.
Vancouver
1. Levy O, Goldberg Y, Dagan I (2015) Improving Distributional Similarity with Lessons Learned from Word Embeddings. In: Transactions of the Association for Computational Linguistics. Association for Computational Linguistics, pp 211–225

BibTeX

@article{levy-etal-2015-improving,
    title = "Improving Distributional Similarity with Lessons Learned from Word Embeddings",
    author = "Levy, Omer  and
      Goldberg, Yoav  and
      Dagan, Ido",
    editor = "Collins, Michael  and
      Lee, Lillian",
    journal = "Transactions of the Association for Computational Linguistics",
    volume = "3",
    year = "2015",
    address = "Cambridge, MA",
    publisher = "MIT Press",
    url = "https://aclanthology.org/Q15-1016/",
    doi = "10.1162/tacl_a_00134",
    pages = "211--225"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/