SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

Felix HillRoi ReichartAnna Korhonen

article2014International Conference on Computational Logic1,355 citations

Introduces SimLex-999, a gold-standard evaluation benchmark that isolates true semantic similarity from conceptual association across diverse parts of speech and concreteness levels, exposing performance gaps in vector space models.

Listen

Natural language processing systems rely on computational models of word meaning for applications such as machine translation, automated dictionary construction, and semantic parsing. However, prevailing evaluation benchmarks confound semantic similarity—whether two concepts share core defining properties, such as a cup and a mug—with topical association, where two concepts merely co-occur frequently, such as coffee and cup. As a result, existing benchmarks often reward models for associating words rather than truly understanding their meaning, even as automated systems have hit the statistical ceiling of human agreement on those older tests.

The article's main objective is to introduce and validate SimLex-999, a new gold standard resource specifically designed to quantify semantic similarity independently from association, and to evaluate how well leading representation-learning models capture genuine similarity across diverse concept types.

To construct this benchmark, the authors selected 999 word pairs representing nouns, verbs, and adjectives spanning the full concrete-to-abstract spectrum. Using Amazon Mechanical Turk, 500 native English speakers scored these pairs solely on semantic similarity under strict quality controls and calibration checks. The authors then benchmarked leading distributional semantic models—including standard vector space co-occurrence models, Singular Value Decomposition dimensionality reduction, and state-of-the-art neural network embedding models—against SimLex-999 as well as legacy datasets like WordSim-353 and MEN.

The findings reveal that state-of-the-art models perform substantially worse on SimLex-999 than on older benchmarks. Model performance on SimLex-999 ranged from correlation scores of 0.098 to 0.446, falling far below the human inter-annotator agreement ceiling of 0.67. This performance gap is primarily driven by strongly associated but dissimilar concepts, which existing models mistakenly score as highly similar due to textual co-occurrence. In architectural comparisons, neural language models outperformed traditional count-based models on abstract concepts and overall similarity, but count-based models performed better on highly concrete concepts. Furthermore, feeding models input structured with syntactic dependency parsing significantly improved similarity estimation (increasing correlation up to 0.446), whereas narrowing context window sizes yielded mixed results.

These results indicate that current language representations are far less mature than earlier evaluation metrics suggested, presenting functional risks for applications requiring strict semantic fidelity. When systems cannot separate topical relatedness from true synonymy, automated tools risk generating incorrect language translations, flawed taxonomy definitions, or improper text replacements. The success of dependency-informed models demonstrates that capturing genuine similarity requires deeper structural and grammatical context rather than simple word adjacency counts.

Organizations developing or deploying language representation models should shift benchmark evaluations to datasets like SimLex-999 to avoid inflated performance estimates. Engineering teams should prioritize incorporating syntactic dependency structures into training pipelines, particularly for relational concepts such as verbs. Because purely text-based models struggle to capture concrete physical concepts, future research should also explore multi-modal architectures that ground language in perceptual data.

While SimLex-999 provides high statistical consistency and clear separation of semantic properties, it evaluates isolated word pairs rather than words situated in extended sentential contexts. Decision-makers can have high confidence in the benchmark’s reliability for measuring baseline conceptual similarity, but they should account for task-specific contextual demands when deploying these models into live applications.

arXiv: 1408.3456
Cover for SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

Abstract

We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness, so that pairs of entities that are associated but not actually similar [Freud, psychology] have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures.

Table of Contents

  • 1 Introduction
  • 2 Design Motivation
  • 2.1 Similarity and Association
  • 2.1.1 Association and similarity in NLP
  • 2.2 Concepts, part-of-speech and concreteness
  • 2.3 Existing gold standard evaluation resources
  • 3 The SimLex-999 Dataset
  • 3.1 Choice of Concepts
  • 3.2 Question Design
  • 3.3 Context-free rating
  • 3.4 Questionnaire structure
  • 3.5 Participants
  • 3.6 Post-processing
  • 4 Analysis of Dataset
  • 4.1 Inter-annotator agreement
  • 4.2 Response validity: Similarity not association
  • 5 Evaluating Models with SimLex-999
  • 5.1 Semantic models
  • 5.2 Results
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — SimLex-999 Dataset Design and Composition

    definition

    SimLex-999 is a gold-standard evaluation benchmark consisting of 999999 English word pairs designed to quantify genuine semantic similarity (e.g., synonymy, near-synonymy, and hypernymy) rather than semantic relatedness or association (e.g., topical or functional co-occurrence).

    The dataset comprises:

    • 666 noun pairs (600600 associated, 6666 unassociated),
    • 222 verb pairs (200200 associated, 2222 unassociated),
    • 111 adjective pairs (100100 associated, 1111 unassociated).

    Candidate words were sampled from the University of South Florida (USF) Free Association Database and required a part-of-speech (POS) tendency ≥75%\ge 75\% towards their designated category based on the British National Corpus. All pairs consist of words within the same POS category.

    Pairs are stratified across four concreteness quadrants (C1,C2,C3,C4C_1, C_2, C_3, C_4) based on human concreteness ratings on a 1–7 scale:

    • C1C_1: First concept concrete (>4>4) with below-median concreteness difference between concepts,
    • C2C_2: First concept concrete (>4>4) with a less concrete second item and above-median concreteness difference,
    • C3C_3: First concept abstract (≤4\le 4) with below-median concreteness difference,
    • C4C_4: First concept abstract (≤4\le 4) with a more concrete second item and above-median concreteness difference.

    Each POS category is equally partitioned across these four concreteness classes.

  2. Knowl 2 — SimLex-999 Annotation and Quality Control Pipeline

    experimental setup

    SimLex-999 similarity ratings were collected from 500 native English speakers via Amazon Mechanical Turk using the following protocol:

    1. Questionnaire Partitioning: The 999 pairs were divided into 10 tranches of 119 pairs each. Each tranche contained 99 unique pairs and a fixed 20-pair consistency set (14 noun, 4 verb, and 2 adjective pairs) used for cross-annotator calibration.
    2. Interface and Grouping: Words were rated within POS-specific groups of 6–7 pairs on an integer slider scale [0,6][0, 6]. For inter-group calibration, the last pair of the preceding group became the first pair of the next group and was re-rated.
    3. Checkpoint Triads: Three checkpoint multiple-choice questions (one per POS) required participants to choose the single genuinely similar pair from three associated options (e.g., distinguishing [bread,toast][\text{bread}, \text{toast}] from [bread,butter][\text{bread}, \text{butter}] and [stale,bread][\text{stale}, \text{bread}]).
    4. Post-Processing and Filtering:
      • Calibration Adjustment: Raters whose consistency-set mean diverged by >1>1 from the overall consistency-set mean had all their non-extremal ratings shifted by ±1\pm 1.
      • Exclusion Criteria: Annotators were discarded if their average pairwise Spearman correlation was >1>1 standard deviation below the population mean, if their removal increased overall agreement more than 50 other raters (worst 10%), or if they failed any checkpoint question.

    A total of 99 annotators were excluded (leaving at least 36 raters per pair). The final scores were averaged across raters and linearly scaled to the interval [0,10][0, 10].

  3. Knowl 3 — Human Inter-Annotator Agreement Across POS and Concreteness in SimLex-999

    empirical result

    The overall inter-annotator agreement on SimLex-999, measured as the average pairwise Spearman rank correlation ρ\rho between raters, is ρ=0.673\rho = 0.673 (compared to ρ=0.61\rho = 0.61 on WordSim-353 using the same metric). Consistency (defined as the inverse of per-pair standard deviation 1/σ1/\sigma) and agreement vary systematically across grammatical and semantic properties:

    • Adjectives: Highest inter-annotator agreement (ρ=0.792\rho = 0.792) and highest consistency (1/σ=0.8921/\sigma = 0.892). This high consistency is attributed to adjectives frequently co-habiting salient, one-dimensional scales (e.g., freezing, cold, warm, hot).
    • Abstract vs. Concrete Concepts: Abstract concept pairs demonstrate higher agreement (ρ=0.703\rho = 0.703, 1/σ=0.8771/\sigma = 0.877) than concrete concept pairs (ρ=0.614\rho = 0.614, 1/σ=0.7521/\sigma = 0.752).
    • Verbs vs. Nouns: Humans rate verb similarity with higher pairwise agreement (ρ=0.717\rho = 0.717, 1/σ=0.7581/\sigma = 0.758) than noun similarity (ρ=0.612\rho = 0.612, 1/σ=0.7811/\sigma = 0.781).
  4. Knowl 4 — Empirical Dissociation of Similarity and Association Across Semantic Datasets

    empirical result

    Re-annotating the 353 word pairs of WordSim-353 using the SimLex-999 similarity-focused instructions (with 100 native English speakers) produced a substantially lower mean rating (4.074.07) than the original WordSim-353 instructions (5.915.91).

    When evaluating against subsets partitioned by semantic relation:

    • On the WS-Sim subset (pairs judged to be similar or totally unrelated), correlation between SimLex-style ratings and original WordSim-353 ratings is high (Spearman ρ=0.73\rho = 0.73).
    • On the WS-Rel subset (pairs judged to be associated but not similar, plus unrelated pairs), correlation between SimLex-style ratings and original WordSim-353 ratings drops sharply to Spearman ρ=0.38\rho = 0.38.

    This validates that standard relatedness instructions reward associated, dissimilar pairs (such as [movie, theater] or [coffee, cup]), whereas genuine similarity guidelines assign them low scores.

  5. Knowl 5 — WordNet Semantic Relations and Their Similarity versus Association Profiles

    data/table

    Analysis of the 382 word pairs in SimLex-999 that also possess explicit lexicographic relations in WordNet demonstrates that similarity and association diverge markedly across lexical semantic relations.

    Semantic Relation SimLex-999 Rating (0–10) USF Free Association (0–10)
    Synonym 7.70 1.57
    1-Hypernym 6.62 1.06
    2-Hypernym 6.19 0.67
    3-Hypernym 5.73 0.48
    4-Hypernym 5.77 0.19
    5-Hypernym 3.76 0.14
    Cohyponym 4.94 0.82
    Meronym 3.93 0.89
    Antonym 1.72 2.99

    Key observations:

    1. Synonyms attain the highest similarity scores (7.707.70) but only moderate association (1.571.57).
    2. Antonyms have the highest free association strength (2.992.99) among all relation types, but low similarity (1.721.72).
    3. Meronyms / Holonyms (part-whole) are relatively strongly associated (0.890.89) but judged low in similarity (3.933.93).
    4. Hypernymy / Hyponymy shows a monotonic decrease in both similarity and association as taxonomy path distance increases from 1 to 5 steps.
  6. Knowl 6 — Comparative Performance of Distributional Semantic Models on SimLex-999 vs. Association Benchmarks

    data/table

    State-of-the-art distributional semantic models, including Neural Language Models (NLMs), Vector Space Models (VSM), and Singular Value Decomposition (SVD), score dramatically lower on SimLex-999 than on association-based benchmarks (WordSim-353 and MEN), remaining far below the human inter-annotator ceiling on SimLex-999 (Spearman ρ=0.67\rho = 0.67).

    Training Corpus Model Architecture WS-353 (ρ\rho) MEN (ρ\rho) SimLex-999 (ρ\rho)
    Wikipedia (≈\approx1B tokens) Huang et al. (2012) 0.623 0.300 0.098
    Wikipedia (≈\approx852M tokens) Collobert Weston (2008) 0.494 0.575 0.268
    Wikipedia (≈\approx1000M tokens) Mikolov et al. (2013a) Skip-gram 0.655 0.699 0.414
    RCV1 (≈\approx150M tokens) Mikolov et al. (2013a) Skip-gram 0.442 0.433 0.282
    RCV1 (≈\approx150M tokens) VSM (PMI-weighted counts) 0.415 0.437 0.194
    RCV1 (≈\approx150M tokens) SVD (300 dimensions) 0.379 0.480 0.233
    Human Agreement Ceiling 0.611 / 0.756 0.680 0.670

    While prediction-based models (Mikolov et al.) and dimensionality reduction (SVD) reach or approach the human ceiling on WordSim-353 and MEN, the best-performing model on SimLex-999 achieves ρ=0.414\rho = 0.414, establishing SimLex-999 as a substantially more demanding test of representation quality.

  7. Knowl 7 — Vulnerability of Distributional Semantic Models to Semantic Association

    empirical result

    When evaluated on the subset of the 333 most highly associated pairs in SimLex-999 (based on USF free association scores), all distributional semantic models exhibit severe performance drops compared to their performance on the entire 999 pairs:

    • Huang et al. (2012): drops from ρ=0.098\rho = 0.098 (all) to ρ=−0.037\rho = -0.037 (most associated 333).
    • Collobert & Weston (2008): drops from ρ=0.270\rho = 0.270 (all) to ρ=0.070\rho = 0.070 (most associated 333).
    • Mikolov et al. (2013a) Skip-gram: drops from ρ=0.414\rho = 0.414 (all) to ρ=0.260\rho = 0.260 (most associated 333).
    • Levy & Goldberg (2014) Dependency Skip-gram: drops from ρ=0.446\rho = 0.446 (all) to ρ=0.347\rho = 0.347 (most associated 333).

    Because distributional models derive semantic vectors from linguistic co-occurrence, they systematically overestimate the similarity of concept pairs that frequently co-occur but are not genuinely similar (e.g., [coffee, cup], [shrink, grow]).

  8. Knowl 8 — Effect of Dependency-Parsed Contexts on Semantic Similarity Across POS

    empirical result

    Replacing linear bag-of-words (BOW) context windows with syntactic dependency contexts (Levy & Goldberg, 2014) in Skip-gram training improves overall similarity modeling and ability to distinguish similarity from association on SimLex-999:

    • Overall SimLex-999: Increases from ρ=0.414\rho = 0.414 (BOW) to ρ=0.446\rho = 0.446 (dependency).
    • Top 333 Associated Pairs: Increases from ρ=0.260\rho = 0.260 (BOW) to ρ=0.347\rho = 0.347 (dependency), showing that dependency information specifically helps suppress association bias.
    • Performance by POS Category:
      • Verbs: Dependency contexts produce the largest relative gain, rising from ρ=0.27\rho = 0.27 to ρ=0.38\rho = 0.38.
      • Nouns: Increases from ρ=0.38\rho = 0.38 to ρ=0.47\rho = 0.47.
      • Adjectives: Small increase from ρ=0.48\rho = 0.48 to ρ=0.50\rho = 0.50.

    The strong effect on verbs supports cognitive and computational theories of verbs as relational concepts whose semantic representation relies heavily on syntactic argument structure.

  9. Knowl 9 — Evaluation of Context Window Size on Word Similarity Modeling

    empirical result

    Evaluating narrow (window size 2) versus wide (window size 10) context windows on RCV1-trained models yields mixed results that challenge the hypothesis that narrower context windows universally improve semantic similarity estimation:

    • Skip-gram (Mikolov et al.):
      • Window 2: ρ=0.282\rho = 0.282 on all pairs, ρ=0.178\rho = 0.178 on top 333 associated pairs.
      • Window 10: ρ=0.266\rho = 0.266 on all pairs, ρ=0.176\rho = 0.176 on top 333 associated pairs.
      • The narrow window provides only a marginal improvement (+0.016+0.016 overall, +0.002+0.002 associated).
    • SVD Model:
      • Window 2: ρ=0.233\rho = 0.233 on all pairs, ρ=0.009\rho = 0.009 on top 333 associated pairs.
      • Window 10: ρ=0.238\rho = 0.238 on all pairs, ρ=0.070\rho = 0.070 on top 333 associated pairs.
      • For SVD, the larger window (10) outperforms the narrow window (2), especially on the highly associated subset (0.0700.070 vs. 0.0090.009).

    The optimal window size depends on the underlying model architecture and concept types rather than adhering to a uniform rule.

  10. Knowl 10 — Differential Performance of Embedding vs. Count Models Across POS and Concreteness

    empirical result

    Evaluating distributional models on SimLex-999 subsets grouped by POS and concreteness reveals architectural trade-offs:

    1. POS Hierarchy Discrepancy:

      • Human agreement is highest on adjectives (ρ=0.792\rho = 0.792), followed by verbs (ρ=0.717\rho = 0.717) and nouns (ρ=0.612\rho = 0.612).
      • All models (Mikolov, VSM, SVD) perform best on adjectives (e.g., Mikolov(2): ρ=0.436\rho = 0.436), but they reverse the human pattern on verbs and nouns, scoring higher on nouns (Mikolov(2): ρ=0.303\rho = 0.303) than on verbs (Mikolov(2): ρ=0.161\rho = 0.161; VSM(2): ρ=0.038\rho = 0.038; SVD(2): ρ=0.085\rho = 0.085).
    2. Abstract vs. Concrete Quartiles:

      • Skip-gram (Mikolov) performs better on abstract concept pairs (ρ=0.306\rho = 0.306 with window 2, ρ=0.309\rho = 0.309 with window 10) than on concrete concept pairs (ρ=0.248\rho = 0.248 with window 2, ρ=0.236\rho = 0.236 with window 10).
      • Count-based models with window 10 (VSM and SVD) perform better on concrete concepts (VSM(10): ρ=0.270\rho = 0.270; SVD(10): ρ=0.325\rho = 0.325) than on abstract concepts (VSM(10): ρ=0.249\rho = 0.249; SVD(10): ρ=0.209\rho = 0.209).

    This indicates that dense neural embeddings are better suited for abstract semantics, whereas co-occurrence count vectors capture concrete referential semantics more effectively.

Coverage note — None was omitted; all key contributions—dataset design, crowdsourcing protocol, human validation analyses, model benchmarks, association tests, dependency analyses, context window experiments, and POS/concreteness breakdowns—are fully represented.

References

  1. 1.Agirre, Eneko, Enrique Alfonseca, Keith Hall, Jana Kravalova, Marius Paşca, and Aitor Soroa. 2009. A study on similarity and relatedness using distributional and Wordnet-based approaches. In Proceedings of NAACL, Boulder, CO.
  2. 2.Alfonseca, Enrique and Suresh Manandhar. 2002. Extending a lexical ontology by a combination of distributional semantics signatures. In G. Schrieber et al., Knowledge Engineering and Knowledge Management: Ontologies and the Semantic Web. Springer, pages 1–7.
  3. 3.Andrews, Mark, Gabriella Vigliocco, and David Vinson. 2009. Integrating experiential and distributional data to learn semantic representations. Psychological Review, 116(3):463.
  4. 4.Bansal, Mohit, Kevin Gimpel, and Karen Livescu. 2014. Tailoring continuous word representations for dependency parsing. In Proceedings of ACL, Baltimore, MD.
  5. 5.Baroni, Marco, Georgiana Dinu, and Germán Kruszewski. 2014. Don’t count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors. In Proceedings of ACL, Baltimore, MD.
  6. 6.Baroni, Marco and Alessandro Lenci. 2010. Distributional memory: A general framework for corpus-based semantics. Computational Linguistics, 36(4):673–721.
  7. 7.Barsalou, Lawrence W., W. Kyle Simmons, Aron K. Barbey, and Christine D. Wilson. 2003. Grounding conceptual knowledge in modality-specific systems. Trends in Cognitive Sciences, 7(2):84–91.
  8. 8.Beltagy, Islam, Katrin Erk, and Raymond Mooney. 2014. Semantic parsing using distributional semantics and probabilistic logic. In ACL 2014 Workshop on Semantic Parsing.
  9. 9.Bengio, Yoshua, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003. A neural probabilistic language model. The Journal of Machine Learning Research, 3:1137–1155.
  10. 10.Bernardi, Raffaella, Georgiana Dinu, Marco Marelli, and Marco Baroni. 2013. A relatedness benchmark to test the role of determiners in compositional distributional semantics. In Proceedings of ACL, Sofia.
  11. 11.Biemann, Chris. 2005. Ontology learning from text: A survey of methods. LDV Forum, 20(2):75–93.
  12. 12.Bird, Steven. 2006. Nltk: the natural language toolkit. In Proceedings of the COLING/ACL on Interactive Presentation sessions, pages 69–72, Sydney.
  13. 13.Bruni, Elia, Gemma Boleda, Marco Baroni, and Nam-Khanh Tran. 2012a. Distributional semantics in technicolor. In Proceedings of ACL, Jeju Island.
  14. 14.Bruni, Elia, Jasper Uijlings, Marco Baroni, and Nicu Sebe. 2012b. Distributional semantics with eyes: Using image analysis to improve computational representations of word meaning. In Proceedings of the 20th ACM International Conference on Multimedia, Nara.
  15. 15.Budanitsky, Alexander and Graeme Hirst. 2006. Evaluating Wordnet-based measures of lexical semantic relatedness. Computational Linguistics, 32(1):13–47.
  16. 16.Cimiano, Philipp, Andreas Hotho, and Steffen Staab. 2005. Learning concept hierarchies from text corpora using formal concept analysis. J. Artif. Intell. Res. (JAIR), 24:305–339.
  17. 17.Collobert, R. and J. Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In International Conference on Machine Learning, ICML, Helsinki.
  18. 18.Cruse, D. Alan. 1986. Lexical semantics. Cambridge University Press.
  19. 19.Cunningham, Hamish. 2005. Information extraction, automatic. Encyclopedia of language and linguistics, pages 665–677.
  20. 20.Fellbaum, Christiane 1998. WordNet. Wiley Online Library.
  21. 21.Finkelstein, Lev, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2001. Placing search in context: The concept revisited. In Proceedings of the 10th International Conference on World Wide Web, pages 406–414, Hong Kong.
  22. 22.Firth, J. R. 1957. Papers in Linguistics 1934–1951. Oxford University Press.
  23. 23.Gentner, Dedre. 1978. On relational meaning: The acquisition of verb meaning. Child Development, pages 988–998.
  24. 24.Gentner, Dedre. 2006. Why verbs are hard to learn. Action meets word: How Children Learn Verbs, pages 544–564.
  25. 25.Gershman, Anatole, Yulia Tsvetkov, Leonid Boytsov, Eric Nyberg, and Chris Dyer. 2014. Metaphor detection with cross-lingual model transfer. In Proceedings of ACL, Baltimore, MD.
  26. 26.Golub, Gene H. and Christian Reinsch. 1970. Singular value decomposition and least squares solutions. Numerische Mathematik, 14(5):403–420.
  27. 27.Griffiths, Thomas L., Mark Steyvers, and Joshua B. Tenenbaum. 2007. Topics in semantic representation. Psychological Review, 114(2):211.
  28. 28.Haghighi, Aria, Percy Liang, Taylor Berg-Kirkpatrick, and Dan Klein. 2008. Learning bilingual lexicons from monolingual corpora. In Proceedings of ACL 2008, Columbus, OH.
  29. 29.Hassan, Samer and Rada Mihalcea. 2011. Semantic relatedness using salient semantic analysis. In AAAI, San Francisco, CA.
  30. 30.Hatzivassiloglou, Vasileios, Judith L. Klavans, Melissa L. Holcombe, Regina Barzilay, Min-Yen Kan, and Kathleen McKeown. 2001. Simfinder: A flexible clustering tool for summarization. In NAACL Workshop on Automatic Summarization, Pittsburgh, PA.
  31. 31.He, Xiaodong, Mei Yang, Jianfeng Gao, Patrick Nguyen, and Robert Moore. 2008. Indirect-HMM-based hypothesis alignment for combining outputs from machine translation systems. In Proceedings of EMNLP, pages 98–107, Edinburgh.
  32. 32.Hill, Felix, Douwe Kiela, and Anna Korhonen. 2013. Concreteness and corpora: A theoretical and practical analysis. CMCL 2013, page 75, Sofia.
  33. 33.Hill, Felix, Anna Korhonen, and Christian Bentz. 2014. A quantitative empirical analysis of the abstract/concrete distinction. Cognitive Science, 38(1):162–177.
  34. 34.Hill, Felix, Roi Reichart, and Anna Korhonen. 2014. Multi-modal models for concrete and abstract concept meaning. Transactions of the Association for Computational Linguistics (TACL), 2:285–296.
  35. 35.Huang, Eric H., Richard Socher, Christopher D. Manning, and Andrew Y. Ng. 2012. Improving word representations via global context and multiple word prototypes. In Proceedings of ACL, pages 873–882, Jeju Island.
  36. 36.Kiela, Douwe and Stephen Clark. 2014. A systematic study of semantic vector space model parameters. In Proceedings of the 2nd Workshop on Continuous Vector Space Models and their Compositionality (CVSC)@ EACL, pages 21–30, Gothenburg.
  37. 37.Kiela, Douwe, Felix Hill, Anna Korhonen, and Stephen Clark. 2014. Improving multi-modal representations using image dispersion: Why less is sometimes more. In Proceedings of ACL, Baltimore, MD.
  38. 38.Landauer, Thomas K. and Susan T. Dumais. 1997. A solution to Plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge. Psychological Review, 104(2):211.
  39. 39.Leech, Geoffrey, Roger Garside, and Michael Bryant. 1994. Claws4: The tagging of the British National Corpus. In Proceedings of COLING, pages 622–628, Kyoto.
  40. 40.Levy, Omer and Yoav Goldberg. 2014. Dependency-based word embeddings. In Proceedings of ACL, volume 2.
  41. 41.Levy, Omer, Steffen Remus, Chris Biemann, and Idol Dagan. 2015. Do supervised distributional methods really learn lexical inference relations? Proceedings of NAACL, Denver, CO.
  42. 42.Lewis, David D., Yiming Yang, Tony G. Rose, and Fan Li. 2004. Rcv1: A new benchmark collection for text categorization research. The Journal of Machine Learning Research, 5:361–397.
  43. 43.Li, Changliang, Bo Xu, Gaowei Wu, Xiuying Wang, Wendong Ge, and Yan Li. 2014. Obtaining better word representations via language transfer. In A. Gelbukh, editor, Computational Linguistics and Intelligent Text Processing. Springer, pages 128–137.
  44. 44.Li, Mu, Yang Zhang, Muhua Zhu, and Ming Zhou. 2006. Exploring distributional similarity based models for query spelling correction. In Proceedings of ALC, pages 1025–1032.
  45. 45.Luong, Minh-Thang, Richard Socher, and Christopher D. Manning. 2013. Better word representations with recursive neural networks for morphology. CoNLL-2013, page 104, Sofia.
  46. 46.Markman, Arthur B. and Edward J. Wisniewski. 1997. Similar and different: The differentiation of basic-level categories. Journal of Experimental Psychology: Learning, Memory, and Cognition, 23(1):54.
  47. 47.Marton, Yuval, Chris Callison-Burch, and Philip Resnik. 2009. Improved statistical machine translation using monolingually-derived paraphrases. In Proceedings of EMNLP, pages 381–390, Edinburgh.
  48. 48.McRae, Ken, Saman Khalkhali, and Mary Hare. 2012. Semantic and associative relations in adolescents and young adults: Examining a tenuous dichotomy. In Valerie F. Reyna, Sandra B. Chapman, Michael R. Dougherty, and Jere Ed Confrey, editors, The Adolescent Brain: Learning, Reasoning, and Decision Making. American Psychological Association, pages 39–66.
  49. 49.Medelyan, Olena, David Milne, Catherine Legg, and Ian H. Witten. 2009. Mining meaning from Wikipedia. International Journal of Human-Computer Studies, 67(9):716–754.
  50. 50.Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. In Proceedings of International Conference of Learning Representations, Scottsdale, AZ.
  51. 51.Mikolov, Tomas, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems, pages 3111–3119, Lake Tahoe, NV.
  52. 52.Navigli, Roberto. 2009. Word sense disambiguation: A survey. ACM Computing Surveys (CSUR), 41(2):10.
  53. 53.Nelson, Douglas L., Cathy L. McEvoy, and Thomas A. Schreiber. 2004. The University of South Florida free association, rhyme, and word fragment norms. Behavior Research Methods, Instruments, & Computers, 36(3):402–407.
  54. 54.Padó, Sebastian, Ulrike Padó, and Katrin Erk. 2007. Flexible, corpus-based modelling of human plausibility judgements. In Proceedings of EMNLP-CoNLL, pages 400–409, Prague.
  55. 55.Paivio, Allan. 1991. Dual coding theory: Retrospect and current status. Canadian Journal of Psychology/Revue canadienne de psychologie, 45(3):255.
  56. 56.Pedersen, Ted, Siddharth Patwardhan, and Jason Michelizzi. 2004. Wordnet:: Similarity: Measuring the relatedness of concepts. In Demonstration Papers at HLT-NAACL 2004, pages 38–41, New York, NY.
  57. 57.Phan, Xuan-Hieu, Le-Minh Nguyen, and Susumu Horiguchi. 2008. Learning to classify short and sparse text & Web with hidden topics from large-scale data collections. In Proceedings of the 17th International Conference on World Wide Web, pages 91–100, Beijing.
  58. 58.Plaut, David C. 1995. Semantic and associative priming in a distributed attractor network. In Proceedings of CogSci, volume 17, pages 37–42, Pittsburgh, PA.
  59. 59.Recchia, Gabriel and Michael N. Jones. 2009. More data trumps smarter algorithms: Comparing pointwise mutual information with latent semantic analysis. Behavior Research Methods, 41(3):647–656.
  60. 60.Reisinger, Joseph and Raymond Mooney. 2010a. A mixture model with sharing for lexical semantics. In Proceedings of EMNLP, pages 1173–1182, Cambridge, MA.
  61. 61.Reisinger, Joseph and Raymond J. Mooney. 2010b. Multi-prototype vector-space models of word meaning. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 109–117, Los Angeles, CA.
  62. 62.Resnik, Philip. 1995. Using information content to evaluate semantic similarity in a taxonomy. In Proceedings of IJCAI.
  63. 63.Resnik, Philip and Jimmy Lin. 2010. 11 evaluations of NLP systems. The handbook of computational linguistics and natural language processing, 57:271.
  64. 64.Rosch, Eleanor, Carol Simpson, and R. Scott Miller. 1976. Structural bases of typicality effects. Journal of Experimental Psychology: Human Perception and Performance, 2(4):491.
  65. 65.Rose, Tony, Mark Stevenson, and Miles Whitehead. 2002. The Reuters corpus volume 1—from yesterday’s news to tomorrow’s language resources. In LREC, volume 2, pages 827–832, Las Palmas.
  66. 66.Rubenstein, Herbert and John B. Goodenough. 1965. Contextual correlates of synonymy. Communications of the ACM, 8(10):627–633.
  67. 67.Silberer, Carina and Mirella Lapata. 2014. Learning grounded meaning representations with autoencoders. In Proceedings of ACL, Sofia.
  68. 68.Sun, Lin, Anna Korhonen, and Yuval Krymolowski. 2008. Verb class discovery from rich syntactic data. In A. Gelbukh, editor, Computational Linguistics and Intelligent Text processing. Springer, pages 16–27.
  69. 69.Turian, Joseph, Lev Ratinov, and Yoshua Bengio. 2010. Word representations: A simple and general method for semi-supervised learning. In Proceedings of ACL, pages 384–394, Uppsala.
  70. 70.Turney, Peter D. 2012. Domain and function: A dual-space model of semantic relations and compositions. Journal of Artificial Intelligence Research (JAIR), 179(44):533–585.
  71. 71.Turney, Peter D. and Patrick Pantel. 2010. From frequency to meaning: Vector space models of semantics. Journal of Artificial Intelligence Research, 37(1):141–188.
  72. 72.Tversky, Amos. 1977. Features of similarity. Psychological Review, 84(4):327.
  73. 73.Wiebe, Janyce. 2000. Learning subjective adjectives from corpora. In AAAI/IAAI, pages 735–740, Austin, TX.
  74. 74.Williams, Gbolahan K. and Sarabjot Singh Anand. 2009. Predicting the polarity strength of adjectives using Wordnet. In ICWSM, San Jose, CA.
  75. 75.Wu, Zhibiao and Martha Palmer. 1994. Verbs, semantics and lexical selection. In Proceedings of ACL, pages 133–138, Las Cruces, NM.
  76. 76.Yong, Chung and Shou King Foo. 1999. A case study on inter-annotator agreement for word sense disambiguation. In Proceedings of the ACL SIGLEX Workshop on Standardizing Lexical Resources (SIGLEX99), College Park, MD.

Citation

MLA
Hill, F., et al. “SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation”. arXiv, 2014, http://arxiv.org/abs/1408.3456v1.
APA
Hill, F., Reichart, R., & Korhonen, A. (2014). SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. arXiv. http://arxiv.org/abs/1408.3456v1
Chicago
Hill, F., R. Reichart, and A. Korhonen. 2014. “SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation”. arXiv. http://arxiv.org/abs/1408.3456v1.
Harvard
Hill, F., Reichart, R. and Korhonen, A. (2014) “SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1408.3456v1.
Vancouver
1. Hill F, Reichart R, Korhonen A (2014) SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. arXiv

BibTeX

@article{hill2014simlex,
  title = {SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation},
  author = {Hill, Felix and Reichart, Roi and Korhonen, Anna},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1408.3456v1},
  eprint = {1408.3456}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF