Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words
Kaitlyn ZhouKawin EthayarajhDallas CardDan Jurafsky
Demonstrates that cosine similarity systematically underestimates the semantic similarity of high-frequency words in contextual embedding models like BERT relative to human judgment due to frequency-induced geometric distortions in embedding space.
Modern natural language processing systems rely heavily on cosine similarity to compare contextual word representations in tasks such as question answering, information retrieval, machine translation, and automated evaluation metrics like BERTScore. However, standard similarity metrics can exhibit hidden biases based on properties of the underlying training data. As language models continue to scale and integrate into high-stakes decision-making pipelines, understanding whether these mathematical similarity measures accurately reflect human understanding is critical.
The main objective of the article is to demonstrate that cosine similarity systematically underestimates the semantic similarity of high-frequency words relative to human judgements and to explain how training data frequency alters the geometric structure of word representations.
To evaluate this effect, the authors conducted statistical regression analyses on contextual embeddings generated by BERT-base-cased across two benchmark datasets containing human similarity annotations: the Word-In-Context dataset and the Stanford Contextualized Word Similarity dataset. Word frequencies were estimated from pre-training corpora encompassing Wikipedia and BookCorpus. In addition, the study sampled nearly 40,000 words to measure the spatial dispersion of contextual representations across high-dimensional space.
The investigation revealed four key findings. First, word frequency in pre-training data is significantly and negatively associated with cosine similarity; frequent words systematically receive lower similarity scores across both identical and differing target words. Second, when predicting human judgements of meaning, cosine similarity exhibits severe underestimation for common words; in the highest frequency tier of the Word-In-Context dataset, cosine similarity predicted that only 25% of pairs shared the same meaning, whereas human annotators judged 54% to be the same. Third, this distortion persists even after controlling for potential confounding factors, including the number of dictionary senses, part of speech, and exact word form. Fourth, geometric analysis showed a strong positive correlation between word frequency and the spatial dispersion of word representations, demonstrating that high-frequency words occupy significantly larger bounding regions in vector space, which mathematically forces a smaller fraction of contextual instances to meet standard similarity thresholds.
These findings imply that standard embedding-based similarity metrics risk misrepresenting common terms, introducing systematic measurement errors into downstream evaluation tools and language applications. This distortion can degrade system performance and introduce socioeconomic disparities, particularly because training data frequency often correlates with geopolitical and demographic representation. The results also show that this issue cannot be fixed by simply applying a simple linear frequency correction.
To mitigate these issues, decision-makers and developers should improve dataset documentation by reporting training frequencies and potential representation distortions to downstream users. Organizations building or evaluating language systems should investigate frequency-aware training interventions and explore geometric post-processing methods to align vector similarities more closely with human perception before relying on standard cosine thresholds.
The conclusions are supported with high statistical confidence across multiple regression models and spatial metrics. However, readers should note that the empirical evaluations focused primarily on BERT-based representations and specific contextual benchmarks. Further validation across alternative model architectures and varying sentence contexts is advisable before implementing large-scale operational modifications.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). Its analysis of contextual embeddings’ anisotropy and context-dependent geometry provides the foundation for understanding why representation dispersion can distort cosine similarity.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). Because the source analyzes BERT representations, BERT’s pretraining and bidirectional contextualization explain the model whose embedding geometry it tests.
No sufficiently relevant recommendations were found.
