Built independently by an author, for readers. Read the story and support ChapterPal

keyword

token representations

Token representations are numerical vector embeddings that represent individual discrete units of data, such as words, subwords, characters, or image patches, within a continuous mathematical space in machine learning models. Initially generated from an input vocabulary through an embedding layer, these vectors are dynamically updated as they pass through the successive hidden layers of neural architectures such as Transformers. Through mechanisms like self-attention and feed-forward networks, token representations aggregate information from surrounding elements, evolving from isolated feature vectors into rich, context-dependent states. These layered representations encode semantic, syntactic, and relational properties across the sequence, allowing the model to perform intermediate computations, route information, and generate final outputs such as probability distributions over a target vocabulary.

6 items

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg

OrganizationsAllen Institute for AIBar-Ilan University

Why you should read this

Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.

Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.

Added

2026-09-29

Layer by Layer: Uncovering Hidden Representations in Language Models

Layer by Layer: Uncovering Hidden Representations in Language Models

Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel, Jalal Naghiyev, Yann LeCun, Ravid Shwartz-Ziv

OrganizationsMetaMila – Québec Artificial Intelligence InstituteNew York UniversityUniversité de MontréalUniversity of California, Los AngelesUniversity of KentuckyWand.AI

Why you should read this

Demonstrates that intermediate layers in language models consistently produce richer representations than the final layer, introducing a geometric and information-theoretic framework that explains why mid-depth embeddings achieve superior performance across diverse downstream tasks.

From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers capture only low-level cues. However, our analysis shows that intermediate layers can encode even richer representations, often improving performance on a range of downstream tasks. To explain and quantify these hidden-layer properties, we propose a unified framework of representation quality metrics based on information theory, geometry, and invariance to input perturbations. Our framework highlights how each layer balances information compression and signal preservation, revealing why mid-depth embeddings can exceed the last layer's performance. Through extensive experiments on 32 text-embedding tasks across various architectures (transformers, state-space models) and domains (language, vision), we demonstrate that intermediate layers consistently provide stronger features, challenging the standard view on final-layer embeddings and opening new directions on using mid-layer representations for more robust and accurate representations.

Added

2026-09-28

GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers

GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers

Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar

OrganizationsIran University of Science and TechnologyKhatam UniversityTehran Institute for Advanced StudiesUniversity of Tehran

Why you should read this

Proposes GlobEnc, a layer-aggregation framework that measures global token attribution by accounting for full Transformer encoder components, including layer normalizations and residual connections, to achieve higher correlation with gradient-based saliency scores than attention-only baselines.

There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies have shown that integrating other components can yield more accurate explanations. This paper introduces a novel token attribution analysis method that incorporates all the components in the encoder block and aggregates this throughout layers. Through extensive quantitative and qualitative experiments, we demonstrate that our method can produce faithful and meaningful global token attributions. Our experiments reveal that incorporating almost every encoder component results in increasingly more accurate analysis in both local (single layer) and global (the whole model) settings. Our global attribution analysis significantly outperforms previous methods on various tasks regarding correlation with gradient-based saliency scores. Our code is freely available at https://github.com/mohsenfayyaz/GlobEnc.

Added

2026-09-26

A Smaller Transformer in Your Transformer

A Smaller Transformer in Your Transformer

Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Adín Ramírez Rivera

OrganizationsDepartment of InformaticsInstitute for Cancer Genetics and InformaticsOslo University HospitalSFI Visual IntelligenceUiT The Arctic University of NorwayUniversity of Oslo

Why you should read this

Proposes Transformer-Within-Transformer, a post-hoc compression technique that fuses contiguous redundant Vision Transformer layers into single learned surrogate blocks, halving model depth and inference computation while matching or exceeding baseline accuracy across natural image and histopathology domains.

Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.

Added

2026-09-20

Creative Commons License
On the Representation Collapse of Sparse Mixture of Experts

On the Representation Collapse of Sparse Mixture of Experts

Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei

OrganizationsBeijing Institute of TechnologyMicrosoftPeking University

Why you should read this

Demonstrates how routing mixture-of-experts models on a low-dimensional hypersphere prevents the representation collapse problem that plagues standard sparse expert systems, resulting in more stable and effective multilingual language models.

Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.

Added

2026-02-23