Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor GevaAvi CaciularuKevin Ro WangYoav Goldberg
Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.
Modern transformer-based language models drive significant advancements in artificial intelligence, yet their internal prediction mechanisms remain largely opaque. Understanding how these systems construct outputs is essential for improving transparency, safety, and operational efficiency as models are deployed across high-stakes industries.
The article aims to evaluate and explain how feed-forward network layers—a fundamental building block in transformer architectures—update internal token representations and build next-token probability distributions in the vocabulary space.
To investigate this, the researchers reverse-engineered feed-forward layers by decomposing their outputs into individual parameter vectors, termed sub-updates, across two representative autoregressive language models: a 16-layer model (WikiLM) and GPT-2. They evaluated these sub-updates by projecting them directly into the output vocabulary space, conducting structured expert annotations to identify semantic and syntactic concept patterns, and analyzing the promotion of candidate tokens using 2,000 validation samples. Additionally, they tested practical applications in toxic content mitigation and computational efficiency using validation sets containing up to 10,000 examples.
The analysis produced several key findings. First, projecting individual parameter sub-updates to the vocabulary revealed human-interpretable concepts—such as "breakfast foods" or "pronouns"—in 36.7% to 55.1% of top-scoring tokens, whereas projecting whole aggregated layer updates obscured these patterns. Second, feed-forward layers operate primarily through a token promotion mechanism rather than token elimination; dominant sub-updates systematically boost favorable candidates (producing maximum positive scores between 1.2 and 8.5) while eliminated candidates receive near-zero mean scores. Third, directly boosting only 10 manually selected safety-related sub-updates in GPT-2 reduced toxic text generation by 47% on a benchmark of challenging toxic prompts, outperforming established self-debiasing methods (37% reduction) and word filtering (20% reduction) with minimal perplexity impact. Fourth, identifying dominant sub-updates enabled an early-exit prediction rule that attained 94.1% accuracy while saving an average of 20% in computational layer processing without retraining the underlying model.
These findings indicate that transformer internal representations can be directly interpreted and steered at the individual vector level rather than treated as uninterpretable black boxes. For organizations deploying language models, this provides a mechanism to mitigate safety and reputational risks through targeted behavioral control while cutting cloud inference costs and energy consumption via self-supervised early exits. It shifts the paradigm from coarse input-level prompt engineering toward precise, internal parameter-level steering.
Organizations should consider piloting vector-level interventions to suppress harmful outputs and exploring early-exit inference strategies to reduce compute costs. However, technical leadership should exercise caution, as toxic language suppression reduces the likelihood of toxic outputs rather than guaranteeing complete elimination, and vector interventions can slightly increase perplexity. Furthermore, because these experiments were conducted specifically on autoregressive decoder models using standard tokenizers, additional analysis is recommended to validate these mechanisms on encoder models, masked language architectures, and diverse domain tasks before production deployment.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This foundational work establishes that transformer feed-forward layers operate as key-value memories storing vocabulary distributions, which directly serves as the basis for reverse-engineering FFN updates in vocabulary space.
- Paper: Dissecting Recall of Factual Associations in Auto-Regressive Language Models, Mor Geva et al. (2023). This work builds directly upon the vocabulary projection and FFN concept-promotion insights to trace the precise circuit by which autoregressive transformers recall and assemble factual associations across layers.
- Paper: Do Llamas Work in English? On the Latent Language of Multilingual Transformers, Chris Wendler et al. (2024). This paper applies vocabulary projection lenses across intermediate transformer layers to analyze the latent conceptual dynamics and multilingual processing stages in large language models.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). This study extends the mechanistic analysis and targeted steering of internal representations to systematically monitor and control high-level model behaviors.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). This paper formalizes the geometric representation of concepts in vocabulary and hidden spaces that underlies the concept promotion and directional steering observed in transformer layers.
