Built independently by an author, for readers. Read the story and support ChapterPal

keyword

dominant sub-updates

Dominant sub-updates are the most influential or highly weighted component vectors within the overall additive update produced by a feed-forward network layer in a transformer model. In transformer architectures, a feed-forward layer calculation can be decomposed into a sum of individual sub-updates, each corresponding to a specific parameter vector scaled by its activation value. Dominant sub-updates are the small subset of these component vectors that receive the strongest activations and contribute the largest magnitude changes to the token representation. By driving the primary changes in the latent state, these dominant components largely govern how the model promotes specific semantic concepts or vocabulary tokens, making them central to analyzing mechanistic predictions, controlling model behaviors, and optimizing computational efficiency.

1 item

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg

OrganizationsAllen Institute for AIBar-Ilan University

Why you should read this

Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.

Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.

Added

2026-09-29