Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

Kaitlyn ZhouKawin EthayarajhDallas CardDan Jurafsky

article2022ACL107 citations

Demonstrates that cosine similarity systematically underestimates the semantic similarity of high-frequency words in contextual embedding models like BERT relative to human judgment due to frequency-induced geometric distortions in embedding space.

Listen

Modern natural language processing systems rely heavily on cosine similarity to compare contextual word representations in tasks such as question answering, information retrieval, machine translation, and automated evaluation metrics like BERTScore. However, standard similarity metrics can exhibit hidden biases based on properties of the underlying training data. As language models continue to scale and integrate into high-stakes decision-making pipelines, understanding whether these mathematical similarity measures accurately reflect human understanding is critical.

The main objective of the article is to demonstrate that cosine similarity systematically underestimates the semantic similarity of high-frequency words relative to human judgements and to explain how training data frequency alters the geometric structure of word representations.

To evaluate this effect, the authors conducted statistical regression analyses on contextual embeddings generated by BERT-base-cased across two benchmark datasets containing human similarity annotations: the Word-In-Context dataset and the Stanford Contextualized Word Similarity dataset. Word frequencies were estimated from pre-training corpora encompassing Wikipedia and BookCorpus. In addition, the study sampled nearly 40,000 words to measure the spatial dispersion of contextual representations across high-dimensional space.

The investigation revealed four key findings. First, word frequency in pre-training data is significantly and negatively associated with cosine similarity; frequent words systematically receive lower similarity scores across both identical and differing target words. Second, when predicting human judgements of meaning, cosine similarity exhibits severe underestimation for common words; in the highest frequency tier of the Word-In-Context dataset, cosine similarity predicted that only 25% of pairs shared the same meaning, whereas human annotators judged 54% to be the same. Third, this distortion persists even after controlling for potential confounding factors, including the number of dictionary senses, part of speech, and exact word form. Fourth, geometric analysis showed a strong positive correlation between word frequency and the spatial dispersion of word representations, demonstrating that high-frequency words occupy significantly larger bounding regions in vector space, which mathematically forces a smaller fraction of contextual instances to meet standard similarity thresholds.

These findings imply that standard embedding-based similarity metrics risk misrepresenting common terms, introducing systematic measurement errors into downstream evaluation tools and language applications. This distortion can degrade system performance and introduce socioeconomic disparities, particularly because training data frequency often correlates with geopolitical and demographic representation. The results also show that this issue cannot be fixed by simply applying a simple linear frequency correction.

To mitigate these issues, decision-makers and developers should improve dataset documentation by reporting training frequencies and potential representation distortions to downstream users. Organizations building or evaluating language systems should investigate frequency-aware training interventions and explore geometric post-processing methods to align vector similarities more closely with human perception before relying on standard cosine thresholds.

The conclusions are supported with high statistical confidence across multiple regression models and spatial metrics. However, readers should note that the empirical evaluations focused primarily on BERT-based representations and specific contextual benchmarks. Further validation across alternative model architectures and varying sentence contexts is advisable before implementing large-scale operational modifications.

No sufficiently relevant recommendations were found.

Cover for Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

Abstract

Cosine similarity of contextual embeddings is used in many NLP tasks (e.g., QA, IR, MT) and metrics (e.g., BERTScore). Here, we uncover systematic ways in which word similarities estimated by cosine over BERT embeddings are understated and trace this effect to training data frequency. We find that relative to human judgements, cosine similarity underestimates the similarity of frequent words with other instances of the same word or other words across contexts, even after controlling for polysemy and other factors. We conjecture that this underestimation of similarity for high frequency words is due to differences in the representational geometry of high and low frequency words and provide a formal argument for the two-dimensional case.

Table of Contents

  • Introduction
  • Effect of Frequency on Cosine Similarity
  • Study 1: WiC
  • Study 2: SCWS
  • Minimum Bounding Hyperspheres
  • Theoretical Intuition
  • Discussion and Conclusion
  • Acknowledgements
  • References
  • A Appendix
  • Regression Results from WiC Experiments
  • Results from SCWS Experiments
  • Results from SCWS Experiments, Explaining the Difference Between Cosine Similarity and Human Judgements
  • Results from Minimum Bounding Hypersphere Experiments
  • Other Ways of Measuring Siblinghood Among Word Pairs
  • Residuals of Predicted Cosine Versus...

Knowls

  1. Knowl 1 — BERT cosine decreases with frequency even within human-labeled sense groups

    empirical result

    On 5,423 Word-in-Context (WiC) examples, the authors represented each target word using BERT-base-cased: they averaged its representations from the last four hidden layers and averaged subword-piece vectors for out-of-vocabulary words. They estimated training frequency from counts in a combination of Wikipedia and BookCorpus data, and modeled cosine similarity between paired contextual embeddings using ordinary least squares (OLS). Within both expert-labeled groups—same meaning (n=2,710n=2{,}710) and different meaning (n=2,713n=2{,}713)—higher log-base-2 frequency predicted lower cosine similarity. In models controlling for log-base-2 WordNet sense count, whether the target wordforms matched, and noun versus verb, the frequency coefficients were −0.0096-0.0096 for same-meaning pairs (R2=0.242R^2=0.242) and −0.0126-0.0126 for different-meaning pairs (R2=0.204R^2=0.204); both frequency effects were significant. Thus, the negative association was not explained merely by the human sense label or these measured confounders.

  2. Knowl 2 — Cosine misses much of the human-judged similarity among the most frequent WiC words

    empirical result

    For WiC, the authors converted cosine similarity to a same- versus different-meaning prediction using a threshold of 0.80.8, tuned on the training set; dev accuracy was 0.660.66. In the highest-frequency decile, humans labeled 54%54\% of examples as having the same meaning, whereas the cosine-based predictions labeled only 25%25\% as same-meaning. The frequency-decile plot on page 3 visualizes this widening disagreement at the high-frequency end. This comparison shows that the downward frequency association in cosine corresponds to underestimation relative to human judgments, not simply to changes in the proportion of human-labeled similar examples.

  3. Knowl 3 — The frequency effect also holds for graded similarity within and across words

    empirical result

    In the Stanford Contextualized Word Similarity dataset (SCWS), crowd workers rated contextual word-pair similarity on a 1–10 scale. The authors separated pairs with the same target word from pairs with different target words and regressed BERT cosine similarity on the average log-base-2 frequency of the two words, human average rating, and average log-base-2 sense count. In the same-target subset (n=214n=214), frequency had coefficient −0.0161-0.0161, human rating +0.0198+0.0198, and sense count −0.0192-0.0192 (R2=0.343R^2=0.343). In the different-target subset (n=1,406n=1{,}406), the respective coefficients were −0.0076-0.0076, +0.0199+0.0199, and −0.0010-0.0010 (R2=0.337R^2=0.337), with the sense-count effect not significant. Frequency therefore remained negatively associated with cosine after accounting for graded human ratings, both for repeated words across contexts and for comparisons between different words.

  4. Knowl 4 — Adding frequency only slightly improves prediction of human similarity ratings

    empirical result

    The authors tested whether frequency could compensate for cosine's underestimation by using SCWS features to predict human ratings. In a full OLS model containing cosine, average log-base-2 frequency, average sense count, and an indicator for whether the target words were the same, the frequency coefficient was positive (+0.0757+0.0757 rating points per unit of average log-base-2 frequency; p=0.006p=0.006), consistent with a correction for underestimation. The model's R2R^2 was 0.4460.446, compared with 0.4430.443 when frequency was excluded. Frequency alone explained almost none of the rating variance (R2=0.002R^2=0.002; coefficient −0.0568-0.0568, p=0.080p=0.080). In these analyses, a linear frequency adjustment provided only a small improvement over cosine and the other features.

  5. Knowl 5 — A word's contextual variation can be measured by its minimum bounding hypersphere

    model/method

    The authors call contextual embeddings of instances of one word type a sibling cohort. They measure the cohort's spatial extent with the radius of its minimum bounding hypersphere: the smallest Euclidean sphere containing all the cohort's embedding vectors. To estimate this quantity, they sampled 10 Wikipedia contexts for each of 39,621 words, generated contextual embeddings with BERT-base-cased, and calculated the radius for each word's 10-vector cohort. This provides a word-level measure of how dispersed contextual representations are, rather than representing each word type by a single point.

  6. Knowl 6 — Higher frequency predicts larger contextual-embedding regions even after accounting for polysemy

    empirical result

    For the 39,621 sampled words, the minimum-bounding-hypersphere radius had a strong positive Pearson correlation with log frequency (r=0.62r=0.62, p<0.001p<0.001). The association also appeared among unique WiC target words (r=0.69r=0.69, p<0.001p<0.001). In OLS models for 1,253 unique WiC words, log-base-2 frequency alone explained 47.7%47.7\% of radius variance and log-base-2 sense count alone explained 44.8%44.8\%; including both explained 58.3%58.3\%, and each predictor remained significant. The frequency–radius relationship is therefore not accounted for solely by the measured number of senses.

  7. Knowl 7 — Larger contextual-embedding regions are associated with lower cosine similarity

    empirical result

    Across 5,412 WiC examples, the minimum-bounding-hypersphere radius of the target word was negatively associated with BERT cosine similarity. An OLS model using radius as its predictor had coefficient −0.0255-0.0255 and R2=0.169R^2=0.169 (p<0.001p<0.001): an increase of one radius unit was associated with a decrease of 0.02550.0255 in cosine similarity. Together with the observed association between frequency and radius, this result connects the broader contextual variation of frequent words with lower cosine scores.

  8. Knowl 8 — Frequency is associated with several other measures of contextual variation

    empirical result

    As a robustness check on the minimum-bounding-hypersphere measure, the authors examined other measures of the space occupied by 10 contextual embeddings per word in a smaller sample. Each measure was positively correlated with log frequency: average pairwise Euclidean distance, r=0.601r=0.601; maximum pairwise Euclidean distance, r=0.584r=0.584; variance of pairwise Euclidean distance, r=0.292r=0.292; average embedding norm, r=0.678r=0.678; and the area of the convex hull after projecting the embeddings to two dimensions with PCA, r=0.603r=0.603. All correlations had p<0.001p<0.001. The frequency–variation pattern is thus not unique to the bounding-hypersphere statistic.

  9. Knowl 9 — A two-dimensional geometric account of cosine underestimation

    theoretical result

    Consider a two-dimensional word vector w\mathbf{w} and a sibling-embedding ball of radius rr centered at vector xc\mathbf{x}_c, with the ball tangent to the line through w\mathbf{w}. Here rr is a nonnegative Euclidean radius and ∥xc∥2\|\mathbf{x}_c\|_2 is the center's distance from the origin, with 0≤r≤∥xc∥20\leq r\leq\|\mathbf{x}_c\|_2. Normalizing points in the ball projects them onto an arc of the unit circle whose angular span is

    2θ=2arcsin⁡(r∥xc∥2).2\theta=2\arcsin\left(\frac{r}{\|\mathbf{x}_c\|_2}\right).

    Cosine between a projected sibling embedding and the normalized comparison vector is their dot product. If human judgments count similarities above a fixed cosine threshold as similar, only part of the arc meets that threshold. Increasing rr while keeping the ball tangent increases the arc span; assuming the distribution of embeddings within the ball does not otherwise change, the fraction meeting the threshold can shrink. This supplies a geometric intuition for why a larger representation region can cause cosine to underestimate similarity relative to human judgments.

  10. Knowl 10 — The geometric account does not explain all observed frequency effects

    limitation

    The two-dimensional geometric argument addresses how the size of a word's embedding region can reduce the share of its instances that reach a cosine threshold against another word. It does not explain why frequent words also have lower cosine similarity to other instances of the same word across contexts; explaining that pattern would require information about how embeddings are distributed within the region. The authors suggest that less anisotropic representations for frequent words may contribute to this within-word effect, but present this as a likely explanation rather than a demonstrated mechanism.

Coverage note — The full appendix regression diagnostics are omitted because the principal coefficients, explained-variance results, and robustness correlations are captured here; no substantial contributed finding is deliberately omitted.

References

  1. 1.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  2. 2.Terra Blevins and Luke Zettlemoyer. 2020. Moving down the long tail of word sense disambiguation with gloss informed bi-encoders. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1006–1017.
  3. 3.Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, page 4356–4364, Red Hook, NY, USA. Curran Associates Inc.
  4. 4.Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota.
  6. 6.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 55–65, Hong Kong, China.
  7. 7.Kawin Ethayarajh, David Duvenaud, and Graeme Hirst. 2019a. Towards understanding linear word analogies. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3253–3262, Florence, Italy.
  8. 8.Kawin Ethayarajh, David Duvenaud, and Graeme Hirst. 2019b. Understanding undesirable word embedding associations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1696–1705, Florence, Italy.
  9. 9.Kawin Ethayarajh and Dan Jurafsky. 2020. Utility is in the eye of the user: A critique of NLP leaderboards. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4846–4853.
  10. 10.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM, 64(12):86–92.
  11. 11.Luke Gessler and Nathan Schneider. 2021. BERT has uncommon sense: Similarity ranking for word sense BERTology. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP.
  12. 12.Nathan Hartmann and Leandro Borges dos Santos. 2018. NILC at CWI 2018: Exploring feature engineering and feature learning. In Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications, pages 335–340, New Orleans, Louisiana.
  13. 13.Johannes Hellrich and Udo Hahn. 2016. Bad Company—Neighborhoods in neural embedding spaces considered harmful. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 2785–2796, Osaka, Japan.
  14. 14.Eric Huang, Richard Socher, Christopher Manning, and Andrew Ng. 2012. Improving word representations via global context and multiple word prototypes. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 873–882, Jeju Island, Korea.
  15. 15.Ting Jiang, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Liangjie Zhang, and Qi Zhang. 2022. Prompt-BERT: Improving BERT sentence embeddings with prompts. arXiv preprint arXiv:2201.04337.
  16. 16.Suyoun Kim, Duc Le, Weiyi Zheng, Tarun Singh, Abhinav Arora, Xiaoyu Zhai, Christian Fuegen, Ozlem Kalinli, and Michael L. Seltzer. 2021. Evaluating user perception of speech recognition system quality with semantic distance metric.
  17. 17.Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9119–9130.
  18. 18.Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela. 2021. Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking. Advances in Neural Information Processing Systems, 34.
  19. 19.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019. Putting evaluation in context: Contextual embeddings improve machine translation evaluation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  20. 20.David Mimno and Laure Thompson. 2017. The strange geometry of skip-gram with negative sampling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2873–2878, Copenhagen, Denmark.
  21. 21.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, page 220–229, New York, NY, USA. Association for Computing Machinery.
  22. 22.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1267–1273, Minneapolis, Minnesota.
  23. 23.Marten Postma, Ruben Izquierdo Bevia, and Piek Vossen. 2016. More is not always better: balancing sense distributions for all-words word sense disambiguation. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 3496–3506, Osaka, Japan. The COLING 2016 Organizing Committee.
  24. 24.Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019. Visualizing and measuring the geometry of bert. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  25. 25.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China.
  26. 26.William Timkey and Marten van Schijndel. 2021. All bark and no bite: Rogue dimensions in transformer language models obscure representational quality. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4527–4546, Online and Punta Cana, Dominican Republic.
  27. 27.Laura Wendlandt, Jonathan K. Kummerfeld, and Rada Mihalcea. 2018. Factors influencing the surprising instability of word embeddings. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2092–2102, New Orleans, Louisiana.
  28. 28.Tianyi Zhang, Varsha Kishore, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
  29. 29.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana.
  30. 30.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China.
  31. 31.Kaitlyn Zhou, Kawin Ethayarajh, and Dan Jurafsky. 2022. Richer countries and richer representations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.
  32. 32.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 19–27.

Citation

MLA
Zhou, K., et al. “Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2022, pp. 401–23, https://doi.org/10.18653/v1/2022.acl-short.45.
APA
Zhou, K., Ethayarajh, K., Card, D., & Jurafsky, D. (2022). Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 401–423. https://doi.org/10.18653/v1/2022.acl-short.45
Chicago
Zhou, K., K. Ethayarajh, D. Card, and D. Jurafsky. 2022. “Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 401–23. https://doi.org/10.18653/v1/2022.acl-short.45.
Harvard
Zhou, K. et al. (2022) “Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 401–423. Available at: https://doi.org/10.18653/v1/2022.acl-short.45.
Vancouver
1. Zhou K, Ethayarajh K, Card D, Jurafsky D (2022) Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 401–423

BibTeX

@inproceedings{zhou-etal-2022-problems,
    title = "Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words",
    author = "Zhou, Kaitlyn  and
      Ethayarajh, Kawin  and
      Card, Dallas  and
      Jurafsky, Dan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-short.45/",
    doi = "10.18653/v1/2022.acl-short.45",
    pages = "401--423"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/