keyword
transformer-based language models
Transformer-based language models are computational systems designed to process, understand, and generate natural language using the transformer neural network architecture. Unlike earlier recurrent designs that analyze text sequentially, these models rely on stacked layers of self-attention mechanisms and feed-forward networks to process entire sequences in parallel, effectively capturing long-range contextual relationships and dependencies among tokens. Typically trained on vast text corpora using self-supervised objectives such as auto-regressive next-token prediction or masked language modeling, these models encode linguistic patterns, semantic structures, and factual associations directly within their parameters. This architecture enables highly parallelized computation and scalable learning, serving as the primary foundation for modern natural language processing tasks including text generation, translation, summarization, and information extraction.
4 items

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
Why you should read this
Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.
Added
2026-09-29

Dissecting Recall of Factual Associations in Auto-Regressive Language Models
Mor Geva, Jasmijn Bastings, Katja Filippova, Amir Globerson
Why you should read this
Reveals the step-by-step internal mechanism auto-regressive language models use to recall facts, demonstrating that early feed-forward layers enrich subject representations while upper attention heads directly extract the correct attributes.
Transformer-based language models (LMs) are known to capture factual knowledge in their parameters. While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference. We investigate this question through the lens of information flow. Given a subject-relation query, we study how the model aggregates information about the subject and relation to predict the correct attribute. With interventions on attention edges, we first identify two critical points where information propagates to the prediction: one from the relation positions followed by another from the subject positions. Next, by analyzing the information at these points, we unveil a three-step internal mechanism for attribute extraction. First, the representation at the last-subject position goes through an enrichment process, driven by the early MLP sublayers, to encode many subject-related attributes. Second, information from the relation propagates to the prediction. Third, the prediction representation “queries” the enriched subject to extract the attribute. Perhaps surprisingly, this extraction is typically done via attention heads, which often encode subject-attribute mappings in their parameters. Overall, our findings introduce a comprehensive view of how factual associations are stored and extracted internally in LMs, facilitating future research on knowledge localization and editing.¹
Added
2026-09-28

Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, David Bau
Why you should read this
Demonstrates through linear probing and causal state transplantation that single intermediate hidden states in autoregressive transformers encode information capable of predicting tokens multiple steps into the future.
Added
2026-09-26

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro
Why you should read this
Demonstrates a novel approach for dynamically allocating compute in transformer-based language models, significantly reducing FLOPs per forward pass and accelerating inference without sacrificing performance, thereby offering a more efficient and hardware-friendly alternative to traditional conditional computation.
Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or compute) to specific positions in a sequence, optimising the allocation along the sequence for different layers across the model depth. Our method enforces a total compute budget by capping the number of tokens () that can participate in the self-attention and MLP computations at a given layer. The tokens to be processed are determined by the network using a top- routing mechanism. Since is defined a priori, this simple procedure uses a static computation graph with known tensor sizes, unlike other conditional computation techniques. Nevertheless, since the identities of the tokens are fluid, this method can expend FLOPs non-uniformly across the time and model depth dimensions. Thus, compute expenditure is entirely predictable in sum total, but dynamic and context-sensitive at the token-level. Not only do models trained in this way learn to dynamically allocate compute, they do so efficiently. These models match baseline performance for equivalent FLOPS and wall-clock times to train, but require a fraction of the FLOPs per forward pass, and can be upwards of 50\% faster to step during post-training sampling.
Added
2026-02-23

