keyword
BERT representations
BERT representations are contextualized vector embeddings generated by the Bidirectional Encoder Representations from Transformers model to represent words, phrases, or entire text sequences as numerical values. Unlike static word embeddings that assign a single fixed vector to each word regardless of usage, BERT representations are dynamically calculated using bidirectional self-attention, allowing the vector for a word to adapt based on surrounding context. Across the architecture of the neural network, these internal hidden-layer vectors capture a hierarchy of linguistic properties, where lower layers typically encode surface and lexical information, intermediate layers capture syntactic structures and grammatical relationships, and deeper layers represent complex semantic meaning and long-range dependencies.
2 items

What Does BERT Learn about the Structure of Language?
Ganesh Jawahar, Benoît Sagot, Djamé Seddah
Why you should read this
Demonstrates how BERT encodes a bottom-to-top hierarchy of surface, syntactic, and semantic information across its layers while implicitly representing tree-like compositional structures to resolve long-distance dependencies.
BERT is a recent language representation model that has surprisingly performed well in diverse language understanding benchmarks. This result indicates the possibility that BERT networks capture structural information about language. In this work, we provide novel support for this claim by performing a series of experiments to unpack the elements of English language structure learned by BERT. We first show that BERT’s phrasal representation captures phrase-level information in the lower layers. We also show that BERT’s intermediate layers encode a rich hierarchy of linguistic information, with surface features at the bottom, syntactic features in the middle and semantic features at the top. BERT turns out to require deeper layers when long-distance dependency information is required, e.g. to track subject-verb agreement. Finally, we show that BERT representations capture linguistic information in a compositional way that mimics classical, tree-like structures.
Added
2026-09-24

A Primer in BERTology: What We Know About How BERT Works
Anna Rogers, Olga Kovaleva, Anna Rumshisky
Why you should read this
Synthesizes over 150 interpretability studies to explain what linguistic knowledge BERT encodes, how its attention mechanisms operate, and how compression techniques address its overparameterization.
Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue and approaches to compression. We then outline directions for future research.
Added
2026-09-18

