keyword
token representations
Token representations are numerical vector embeddings that represent individual discrete units of data, such as words, subwords, characters, or image patches, within a continuous mathematical space in machine learning models. Initially generated from an input vocabulary through an embedding layer, these vectors are dynamically updated as they pass through the successive hidden layers of neural architectures such as Transformers. Through mechanisms like self-attention and feed-forward networks, token representations aggregate information from surrounding elements, evolving from isolated feature vectors into rich, context-dependent states. These layered representations encode semantic, syntactic, and relational properties across the sequence, allowing the model to perform intermediate computations, route information, and generate final outputs such as probability distributions over a target vocabulary.
6 items

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
Why you should read this
Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.
Added
2026-09-29

Layer by Layer: Uncovering Hidden Representations in Language Models
Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel, Jalal Naghiyev, Yann LeCun, Ravid Shwartz-Ziv
Why you should read this
Demonstrates that intermediate layers in language models consistently produce richer representations than the final layer, introducing a geometric and information-theoretic framework that explains why mid-depth embeddings achieve superior performance across diverse downstream tasks.
From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers capture only low-level cues. However, our analysis shows that intermediate layers can encode even richer representations, often improving performance on a range of downstream tasks. To explain and quantify these hidden-layer properties, we propose a unified framework of representation quality metrics based on information theory, geometry, and invariance to input perturbations. Our framework highlights how each layer balances information compression and signal preservation, revealing why mid-depth embeddings can exceed the last layer's performance. Through extensive experiments on 32 text-embedding tasks across various architectures (transformers, state-space models) and domains (language, vision), we demonstrate that intermediate layers consistently provide stronger features, challenging the standard view on final-layer embeddings and opening new directions on using mid-layer representations for more robust and accurate representations.
Added
2026-09-28

GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
Why you should read this
Proposes GlobEnc, a layer-aggregation framework that measures global token attribution by accounting for full Transformer encoder components, including layer normalizations and residual connections, to achieve higher correlation with gradient-based saliency scores than attention-only baselines.
There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies have shown that integrating other components can yield more accurate explanations. This paper introduces a novel token attribution analysis method that incorporates all the components in the encoder block and aggregates this throughout layers. Through extensive quantitative and qualitative experiments, we demonstrate that our method can produce faithful and meaningful global token attributions. Our experiments reveal that incorporating almost every encoder component results in increasingly more accurate analysis in both local (single layer) and global (the whole model) settings. Our global attribution analysis significantly outperforms previous methods on various tasks regarding correlation with gradient-based saliency scores. Our code is freely available at https://github.com/mohsenfayyaz/GlobEnc.
Added
2026-09-26

Measuring the Mixing of Contextual Information in the Transformer
Javier Ferrando, Gerard I. Gállego, Marta R. Costa-jussà
Why you should read this
Proposes ALTI, an interpretability method that tracks information flow across full Transformer attention blocks to generate input attributions that surpass gradient-based techniques in faithfulness and reliability.
The Transformer architecture aggregates input information through the self-attention mechanism, but there is no clear understanding of how this information is mixed across the entire model. Additionally, recent works have demonstrated that attention weights alone are not enough to describe the flow of information. In this paper, we consider the whole attention block –multi-head attention, residual connection, and layer normalization– and define a metric to measure token-to-token interactions within each layer. Then, we aggregate layer-wise interpretations to provide input attribution scores for model predictions. Experimentally, we show that our method, ALTI (Aggregation of Layer-wise Token-to-token Interactions), provides more faithful explanations and increased robustness than gradient-based methods.
Added
2026-09-26

A Smaller Transformer in Your Transformer
Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Adín Ramírez Rivera
Why you should read this
Proposes Transformer-Within-Transformer, a post-hoc compression technique that fuses contiguous redundant Vision Transformer layers into single learned surrogate blocks, halving model depth and inference computation while matching or exceeding baseline accuracy across natural image and histopathology domains.
Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.
Added
2026-09-20


On the Representation Collapse of Sparse Mixture of Experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei
Why you should read this
Demonstrates how routing mixture-of-experts models on a low-dimensional hypersphere prevents the representation collapse problem that plagues standard sparse expert systems, resulting in more stable and effective multilingual language models.
Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.
Added
2026-02-23
