Built independently by an author, for readers. Read the story and support ChapterPal

keyword

BERT representations

BERT representations are contextualized vector embeddings generated by the Bidirectional Encoder Representations from Transformers model to represent words, phrases, or entire text sequences as numerical values. Unlike static word embeddings that assign a single fixed vector to each word regardless of usage, BERT representations are dynamically calculated using bidirectional self-attention, allowing the vector for a word to adapt based on surrounding context. Across the architecture of the neural network, these internal hidden-layer vectors capture a hierarchy of linguistic properties, where lower layers typically encode surface and lexical information, intermediate layers capture syntactic structures and grammatical relationships, and deeper layers represent complex semantic meaning and long-range dependencies.

2 items