Built independently by an author, for readers. Read the story and support ChapterPal

keyword

contextuality measures

Contextuality measures are quantitative metrics used in natural language processing to evaluate the degree to which the numerical representation of a word changes depending on its surrounding text. In neural language modeling and representation learning, these measures determine whether a model creates dynamic, context-specific representations for words across different sentences rather than relying on fixed or static embeddings. Typical contextuality measures analyze the geometric properties of high-dimensional vector spaces by calculating metrics such as self-similarity, which compares representations of the same word across diverse linguistic contexts, intra-sentence similarity, which evaluates how representations of different words shift when appearing in the same passage, and explainable variance, which quantifies how much of a contextualized representation can be accounted for by a single static baseline vector. These metrics provide insight into how deeply neural networks encode syntax, semantic nuances, and polysemy across different processing layers.

1 item

How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings

How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings

Kawin Ethayarajh

OrganizationsStanford University

Why you should read this

Reveals the geometric behavior of BERT, ELMo, and GPT-2 embeddings across layers, proving that upper layers generate more context-specific representations while static word vectors account for less than five percent of their variance.

Replacing static word embeddings with contextualized word representations has yielded significant improvements on many NLP tasks. However, just how contextual are the contextualized representations produced by models such as ELMo and BERT? Are there infinitely many context-specific representations for each word, or are words essentially assigned one of a finite number of word-sense representations? For one, we find that the contextualized representations of all words are not isotropic in any layer of the contextualizing model. While representations of the same word in different contexts still have a greater cosine similarity than those of two different words, this self-similarity is much lower in upper layers. This suggests that upper layers of contextualizing models produce more context-specific representations, much like how upper layers of LSTMs produce more task-specific representations. In all layers of ELMo, BERT, and GPT-2, on average, less than 5% of the variance in a word's contextualized representations can be explained by a static embedding for that word, providing some justification for the success of contextualized representations.

Added

2026-09-25