Built independently by an author, for readers. Read the story and support ChapterPal

keyword

vocabulary space

Vocabulary space refers to the high-dimensional mathematical space in which each coordinate dimension corresponds to a distinct token within a language model's predefined vocabulary. In machine learning and neural network interpretability, vectors residing in this space typically represent numerical scores, unnormalized logits, or probability distributions assigned across every possible output token. While deep neural networks like transformers internally manipulate information within continuous, lower-dimensional latent embedding spaces, projecting intermediate activations or component updates into vocabulary space translates abstract internal representations into human-interpretable distributions. This transformation allows researchers to observe, analyze, and manipulate how individual layers and attention heads promote, suppress, or refine specific word predictions and conceptual associations throughout the generation process.

2 items

Characterizing Mechanisms for Factual Recall in Language Models

Characterizing Mechanisms for Factual Recall in Language Models

Qinan Yu, Jack Merullo, Ellie Pavlick

OrganizationsBrown UniversityDepartment of Computer Science

Why you should read this

Demonstrates that competition between memorized facts and counterfactual context in language models is governed by pretraining frequencies and can be dynamically controlled at runtime by scaling individual attention heads.

Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context. These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict. On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations. Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., “The capital of Poland is London”) to overwrite what it learned in pretraining (“Warsaw”). On Pythia and GPT2, the training frequency of both the query country (“Poland”) and the in-context city (“London”) highly affect the models’ likelihood of using the counterfactual. We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits. By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data. This method can increase the rate of generating the in-context answer to 88% of the time simply by scaling a single head at runtime. Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.

Added

2026-10-03

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg

OrganizationsAllen Institute for AIBar-Ilan University

Why you should read this

Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.

Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.

Added

2026-09-29