Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Mor GevaAvi CaciularuKevin Ro WangYoav Goldberg

article2022EMNLP657 citations

Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.

Listen

Modern transformer-based language models drive significant advancements in artificial intelligence, yet their internal prediction mechanisms remain largely opaque. Understanding how these systems construct outputs is essential for improving transparency, safety, and operational efficiency as models are deployed across high-stakes industries.

The article aims to evaluate and explain how feed-forward network layers—a fundamental building block in transformer architectures—update internal token representations and build next-token probability distributions in the vocabulary space.

To investigate this, the researchers reverse-engineered feed-forward layers by decomposing their outputs into individual parameter vectors, termed sub-updates, across two representative autoregressive language models: a 16-layer model (WikiLM) and GPT-2. They evaluated these sub-updates by projecting them directly into the output vocabulary space, conducting structured expert annotations to identify semantic and syntactic concept patterns, and analyzing the promotion of candidate tokens using 2,000 validation samples. Additionally, they tested practical applications in toxic content mitigation and computational efficiency using validation sets containing up to 10,000 examples.

The analysis produced several key findings. First, projecting individual parameter sub-updates to the vocabulary revealed human-interpretable concepts—such as "breakfast foods" or "pronouns"—in 36.7% to 55.1% of top-scoring tokens, whereas projecting whole aggregated layer updates obscured these patterns. Second, feed-forward layers operate primarily through a token promotion mechanism rather than token elimination; dominant sub-updates systematically boost favorable candidates (producing maximum positive scores between 1.2 and 8.5) while eliminated candidates receive near-zero mean scores. Third, directly boosting only 10 manually selected safety-related sub-updates in GPT-2 reduced toxic text generation by 47% on a benchmark of challenging toxic prompts, outperforming established self-debiasing methods (37% reduction) and word filtering (20% reduction) with minimal perplexity impact. Fourth, identifying dominant sub-updates enabled an early-exit prediction rule that attained 94.1% accuracy while saving an average of 20% in computational layer processing without retraining the underlying model.

These findings indicate that transformer internal representations can be directly interpreted and steered at the individual vector level rather than treated as uninterpretable black boxes. For organizations deploying language models, this provides a mechanism to mitigate safety and reputational risks through targeted behavioral control while cutting cloud inference costs and energy consumption via self-supervised early exits. It shifts the paradigm from coarse input-level prompt engineering toward precise, internal parameter-level steering.

Organizations should consider piloting vector-level interventions to suppress harmful outputs and exploring early-exit inference strategies to reduce compute costs. However, technical leadership should exercise caution, as toxic language suppression reduces the likelihood of toxic outputs rather than guaranteeing complete elimination, and vector interventions can slightly increase perplexity. Furthermore, because these experiments were conducted specifically on autoregressive decoder models using standard tokenizers, additional analysis is recommended to validate these mechanisms on encoder models, masked language architectures, and diverse domain tasks before production deployment.

arXiv: 2203.14680aviclu/ffn-values
  • Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This foundational work establishes that transformer feed-forward layers operate as key-value memories storing vocabulary distributions, which directly serves as the basis for reverse-engineering FFN updates in vocabulary space.
Cover for Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Abstract

Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.

Table of Contents

  • 1 Introduction
  • 2 Token Representations as Evolving Distributions Over the Vocabulary
  • 3 The FFN Output as a Collection of Updates to the Output Distribution
  • 4 Sub-Updates Encode Concepts in the Vocabulary Space
  • 4.1 Projection of Sub-Updates is Meaningful
  • 4.2 Sub-Update Projections are Interpretable
  • 5 FFN Updates Promote Tokens in the Output Distribution
  • 5.1 Promoted Versus Eliminated Candidates
  • 5.2 Sub-Updates Across Layers
  • 6 Applications
  • 6.1 Zero-Shot Toxic Language Suppression
  • 6.2 Self-Supervised Early Exit Prediction
  • 7 Related Work
  • 8 Conclusions
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Value Vectors Projection Method
  • A.2 Concepts Annotation
  • A.3 Sub-Update Contribution in FFN Outputs
  • A.4 Toxic Language Suppression Details
  • A.5 Early Exit Details

Knowls

  1. Knowl 1 — Decomposition of Transformer Feed-Forward Updates into Vocabulary-Space Sub-Updates

    equation

    In transformer-based language models, a feed-forward network (FFN) layer ℓ∈{1,…,L}\ell \in \{1, \dots, L\} maps an input contextual token representation xℓ∈Rdx^\ell \in \mathbb{R}^d through key parameter matrix WKℓ∈Rdm×dW_K^\ell \in \mathbb{R}^{d_m \times d}, value parameter matrix WVℓ∈Rdm×dW_V^\ell \in \mathbb{R}^{d_m \times d}, and an activation function ff:

    oℓ=FFNℓ(xℓ)=f(WKℓxℓ)WVℓ=∑i=1dmf(xℓ⋅kiℓ)viℓ=∑i=1dmmiℓviℓo^\ell = \text{FFN}^\ell(x^\ell) = f(W_K^\ell x^\ell) W_V^\ell = \sum_{i=1}^{d_m} f(x^\ell \cdot k_i^\ell) v_i^\ell = \sum_{i=1}^{d_m} m_i^\ell v_i^\ell

    where kiℓ∈Rdk_i^\ell \in \mathbb{R}^d is the ii-th row of WKℓW_K^\ell (the key vector), viℓ∈Rdv_i^\ell \in \mathbb{R}^d is the ii-th column of WVℓW_V^\ell (the value vector), and miℓ=f(xℓ⋅kiℓ)∈Rm_i^\ell = f(x^\ell \cdot k_i^\ell) \in \mathbb{R} is the dynamic scalar activation coefficient.

    The output oℓo^\ell is added to the residual stream to yield updated representation x~ℓ=xℓ+oℓ\tilde{x}^\ell = x^\ell + o^\ell. Projecting into the vocabulary V\mathcal{V} via the token embedding matrix E∈R∣V∣×dE \in \mathbb{R}^{|\mathcal{V}| \times d} produces an additive logit update:

    Ex~ℓ=Exℓ+Eoℓ=Exℓ+∑i=1dmmiℓEviℓE \tilde{x}^\ell = E x^\ell + E o^\ell = E x^\ell + \sum_{i=1}^{d_m} m_i^\ell E v_i^\ell

    The isolated effect of a single sub-update miℓviℓm_i^\ell v_i^\ell on the predicted probability of a token w∈Vw \in \mathcal{V} with embedding vector ew∈Rde_w \in \mathbb{R}^d is:

    p(w∣xℓ+miℓviℓ,E)=exp⁡(ew⋅xℓ+ew⋅miℓviℓ)Z(E(xℓ+miℓviℓ))∝exp⁡(ew⋅xℓ)⋅exp⁡(ew⋅miℓviℓ)p(w \mid x^\ell + m_i^\ell v_i^\ell, E) = \frac{\exp\left(e_w \cdot x^\ell + e_w \cdot m_i^\ell v_i^\ell\right)}{Z(E(x^\ell + m_i^\ell v_i^\ell))} \propto \exp\left(e_w \cdot x^\ell\right) \cdot \exp\left(e_w \cdot m_i^\ell v_i^\ell\right)

    where Z(⋅)Z(\cdot) is the softmax normalizer. Each sub-update thus consists of:

    1. A static scoring vector riℓ=Eviℓ∈R∣V∣r_i^\ell = E v_i^\ell \in \mathbb{R}^{|\mathcal{V}|}, which ranks vocabulary tokens independently of the input context.
    2. A dynamic scalar coefficient miℓm_i^\ell, which scales the vector's contribution for a specific input.
  2. Knowl 2 — Concept Encoding in Transformer Value Vector Projections

    empirical result

    Projecting individual feed-forward value vectors viℓ∈Rdv_i^\ell \in \mathbb{R}^d to the output vocabulary V\mathcal{V} via riℓ=Eviℓr_i^\ell = E v_i^\ell produces rankings whose top-scoring tokens reflect human-interpretable semantic and syntactic concepts, whereas projecting the aggregated FFN layer output EoℓE o^\ell or random vectors yields significantly less interpretable patterns.

    In expert annotations evaluating the top-30 scoring tokens across layers of WikiLM (a 16-layer word-level LM with vocabulary size ∣V∣=267,744|\mathcal{V}| = 267,744) and GPT-2 (a 12-layer subword LM with ∣V∣=50,257|\mathcal{V}| = 50,257):

    • In WikiLM, 55.1% of top tokens across sampled value vectors corresponded to distinct semantic, syntactic, or named concepts (e.g., pronouns, adverbs, groups of people), averaging 1.5 concepts per vector.
    • In GPT-2, 36.7% of top tokens across sampled value vectors matched coherent concepts (e.g., measurement units, WH-relativizers, food/drinks), averaging 1.1 concepts per vector.
    • Random vectors matching the empirical mean and standard deviation of real vectors contained only 22.7% concept-like tokens in WikiLM and 16.0% in GPT-2 (with 0% semantic or syntactic concepts in WikiLM and only 4% semantic in GPT-2).
    • Projecting the aggregated FFN update vector EoℓE o^\ell without decomposition resulted in low non-stopword concept coverage (19.7% in WikiLM and 11.8% in GPT-2) because whole layer updates are dominated by punctuation and stopwords (which comprise 19.7% of WikiLM and 34.2% of GPT-2 update tokens).
  3. Knowl 3 — Token Promotion as the Primary Mechanism of Feed-Forward Updates

    empirical result

    Transformer feed-forward network (FFN) layers shape output probability distributions primarily by promoting target candidate tokens rather than actively suppressing competing candidates.

    In an evaluation of 2,000 validation sequences from WikiText-103 using GPT-2 (L=24,d=1024,dm=4096L=24, d=1024, d_m=4096) and WikiLM (L=16,d=1024,dm=4096L=16, d=1024, d_m=4096):

    • In saturation events—layer updates pℓ→p~ℓp^\ell \to \tilde{p}^\ell where the token that ultimately becomes the final prediction w=arg⁡max⁡(y)w = \arg\max(y) is promoted to rank 1—the 10 most dominant sub-updates (ranked by ∣miℓ∣⋅∥viℓ∥2|m_i^\ell| \cdot \|v_i^\ell\|_2) assign high positive scores to the target token. The average maximum score assigned across events is 8.58.5 in GPT-2 and 1.21.2 in WikiLM, with non-negative mean scores (1.31.3 in GPT-2, <0.01<0.01 in WikiLM).
    • In elimination events—layer updates where the top candidate experiences the largest drop in rank—the top-10 dominant sub-updates assign near-zero mean scores (−0.01-0.01 in WikiLM, 0.10.1 in GPT-2), with much smaller maximum scores (4.04.0 in GPT-2, 0.50.5 in WikiLM).
    • The score distributions for eliminated candidates are symmetric around zero (±0.5\pm 0.5 in WikiLM; +4.0+4.0 vs. −3.6-3.6 in GPT-2), indicating that candidates drop in rank because competing candidates receive positive promotion from dominant sub-updates, rather than being explicitly eliminated by negative scores.
  4. Knowl 4 — Relative Contribution Metric for Feed-Forward Sub-Updates

    definition

    For a feed-forward layer ℓ\ell with intermediate dimension dmd_m, the relative contribution of the ii-th sub-update miℓviℓm_i^\ell v_i^\ell to the layer's output representation is defined as:

    contrib(miℓviℓ):=∣miℓ∣⋅∥viℓ∥2∑j=1dm∣mjℓ∣⋅∥vjℓ∥2\text{contrib}(m_i^\ell v_i^\ell) := \frac{|m_i^\ell| \cdot \|v_i^\ell\|_2}{\sum_{j=1}^{d_m} |m_j^\ell| \cdot \|v_j^\ell\|_2}

    where miℓ∈Rm_i^\ell \in \mathbb{R} is the dynamic activation coefficient, viℓ∈Rdv_i^\ell \in \mathbb{R}^d is the value parameter vector, and ∥⋅∥2\|\cdot\|_2 is the Euclidean norm. Absolute values ∣miℓ∣|m_i^\ell| account for non-monotonic activation functions such as GeLU that produce negative activations.

    Across GPT-2 and WikiLM (dm=4096d_m=4096), the 10 sub-updates with the largest contrib(miℓviℓ)\text{contrib}(m_i^\ell v_i^\ell) account for 5% to over 10% of the total layer output norm sum despite constituting only 10/4096≈0.24%10 / 4096 \approx 0.24\% of the value vectors in that layer.

  5. Knowl 5 — Functional Clusters of Extreme Sub-Updates and Saturation Vectors

    empirical result

    Agglomerative hierarchical clustering of FFN value vectors across all layers using cosine distance D(ℓ1,i1),(ℓ2,i2)=1−cos⁡(vi1ℓ1,vi2ℓ2)D_{(\ell_1, i_1), (\ell_2, i_2)} = 1 - \cos(v_{i_1}^{\ell_1}, v_{i_2}^{\ell_2}) with k=10,000k = 10,000 clusters reveals two specialized functional groups responsible for extreme sub-update activations (top-candidate scores exceeding ±10\pm 10 in GPT-2 or ±5\pm 5 in WikiLM):

    1. Common-Token Vectors: Distributed across all network layers, these vectors promote high-frequency tokens (such as punctuation and stopwords). They activate primarily on short contexts (≤3\le 3 tokens) or trivially predictable tokens (such as sentence-ending periods).
    2. Saturation Vectors: Concentrated in the final layers (layers 23–24 in 24-layer GPT-2 and layers 14–16 in 16-layer WikiLM), these vectors assign large positive weights to generally rare and unlikely tokens. Rather than changing top-candidate rankings, they propagate the residual stream forward without altering top tokens, activating when the model has already finalized its prediction in earlier layers.

    Together, these functional vectors comprise only 1.7% of all value vectors in GPT-2 and 1.1% in WikiLM.

  6. Knowl 6 — Zero-Shot Controlled Generation via FFN Value Vector Steering

    algorithm

    Zero-shot toxic language suppression modifies language model generation by statically overriding the activation coefficients of a selected subset of non-toxic concept value vectors in feed-forward layers.

    Input: Language model with FFN layers ℓ∈{1,…,L}\ell \in \{1, \dots, L\}, prompt sequence w=⟨w1,…,wt⟩w = \langle w_1, \dots, w_t \rangle, set of selected non-toxic value vector indices S={(ℓ1,i1),…,(ℓk,ik)}S = \{(\ell_1, i_1), \dots, (\ell_k, i_k)\}, constant boost coefficient α>0\alpha > 0, generation length TT
    Output: Generated token sequence wt+1,…,wt+Tw_{t+1}, \dots, w_{t+T}
    for step t′=t+1t' = t+1 to t+Tt+T do
        for each layer ℓ=1\ell = 1 to LL do
            Compute intermediate activations mℓ=f(WKℓxℓ)∈Rdmm^\ell = f(W_K^\ell x^\ell) \in \mathbb{R}^{d_m}
            for each (ℓ′,i)∈S(\ell', i) \in S such that ℓ′=ℓ\ell' = \ell do
                miℓ←αm_i^\ell \leftarrow \alpha
            end for
            Compute FFN output oℓ=∑i=1dmmiℓviℓo^\ell = \sum_{i=1}^{d_m} m_i^\ell v_i^\ell
            Update representation x~ℓ=xℓ+oℓ\tilde{x}^\ell = x^\ell + o^\ell
            Pass x~ℓ\tilde{x}^\ell through remaining layer operations
        end for
        Sample next token wt′∼softmax(Ex~L)w_{t'} \sim \text{softmax}(E \tilde{x}^L)
        Append wt′w_{t'} to input sequence
    end for
    return wt+1,…,wt+Tw_{t+1}, \dots, w_{t+T}

    In GPT-2 experiments, k=10k=10 value vectors (0.01% of all vectors) promoting safety and positive concepts (e.g., tokens like "safe", "peace", "modesty", "respectful", "thank") were identified either manually via vocabulary projections in under 5 minutes or automatically via Perspective API toxicity filtering (<0.1<0.1). Setting α=3\alpha = 3 steers generation toward non-toxic text.

  7. Knowl 7 — Toxicity Suppression Performance Comparison on RealToxicityPrompts

    data/table

    Activating 10 safety-aligned FFN value vectors with coefficient α=3\alpha=3 in GPT-2 substantially decreases toxicity on the 1,225 challenging prompts of RealToxicityPrompts (measured across 20-token generations evaluated by the Perspective API with toxicity threshold >0.5> 0.5).

    Model Toxicity Severe Sexually Threat Profanity Identity PPL
    Toxicity Explicit Attack
    GPT-2 (Base) 58.5% 49.2% 34.1% 16.4% 52.5% 16.8% 21.7
    ↑\uparrow 10 Manual Pick 30.8% (↓\downarrow47%) 24.8% (↓\downarrow50%) 20.4% (↓\downarrow40%) 6.0% (↓\downarrow63%) 27.9% (↓\downarrow47%) 8.8% (↓\downarrow48%) 25.3
    ↑\uparrow 10 API Graded 52.7% (↓\downarrow10%) 44.0% (↓\downarrow11%) 33.2% (↓\downarrow3%) 13.3% (↓\downarrow19%) 47.6% (↓\downarrow9%) 15.3% (↓\downarrow9%) 23.8
    Self-Debiasing (SD) 37.2% (↓\downarrow37%) 26.4% (↓\downarrow46%) 21.7% (↓\downarrow36%) 7.8% (↓\downarrow52%) 32.0% (↓\downarrow39%) 8.4% (↓\downarrow50%) 23.9
    WordFilter 46.9% (↓\downarrow20%) 32.4% (↓\downarrow34%) 21.9% (↓\downarrow36%) 16.3% (↓<1\downarrow<1%) 32.3% (↓\downarrow38%) 14.7% (↓\downarrow13%) -

    Steering the 10 manually identified safety-promoting value vectors reduces overall toxicity by 47% relative to base GPT-2, outperforming both the Self-Debiasing baseline (37% reduction, decay constant λ=50\lambda=50) and WordFilter (20% reduction), while causing only a modest increase in language model perplexity (from 21.7 to 25.3).

  8. Knowl 8 — Self-Supervised Early Exiting via Dominant Sub-Update Clustering

    algorithm

    An early exit mechanism determines whether a transformer model can halt inference at intermediate layer ℓ\ell by checking the intersection of active dominant sub-update clusters against pre-collected saturation cluster profiles, without requiring external neural classifier training.

    Input: Pre-clustered value vectors {C1,…,CK}\{C_1, \dots, C_K\}, validation dataset with recorded saturation layer labels, input sequence ww, candidate layer ℓ∈{1,…,L−1}\ell \in \{1, \dots, L-1\}
    Output: Decision whether to exit computation at layer ℓ\ell
    Construct from validation data:
        Tℓ←T^\ell \leftarrow set of top-10 dominant sub-update cluster sets for examples that saturate at layer ℓ\ell
        Nℓ′←N^{\ell'} \leftarrow set of top-10 dominant sub-update cluster sets for examples that saturate at layer ℓ′>ℓ\ell' > \ell
    During inference at layer ℓ\ell:
        Extract top-10 dominant sub-updates: Dℓ=top-10i∣miℓ∣⋅∥viℓ∥2D^\ell = \text{top-10}_{i} |m_i^\ell| \cdot \|v_i^\ell\|_2
        Map DℓD^\ell to cluster IDs: Aℓ={cluster(viℓ)∣i∈Dℓ}A^\ell = \{\text{cluster}(v_i^\ell) \mid i \in D^\ell\}
        Compute overlap metrics with saturation profiles:
            Isat←average and maximum ∣Aℓ∩t∣ for t∈TℓI_{\text{sat}} \leftarrow \text{average and maximum } |A^\ell \cap t| \text{ for } t \in T^\ell
            Inon-sat(ℓ′)←average and maximum ∣Aℓ∩n∣ for n∈Nℓ′I_{\text{non-sat}}(\ell') \leftarrow \text{average and maximum } |A^\ell \cap n| \text{ for } n \in N^{\ell'}
        if Isat>Inon-sat(ℓ′)I_{\text{sat}} > I_{\text{non-sat}}(\ell') for all ℓ′>ℓ\ell' > \ell then
            Halt inference and return prediction arg⁡max⁡wpwℓ\arg\max_w p^\ell_w
        else
            Proceed to layer ℓ+1\ell+1
        end if

    This procedure leverages the characteristic cluster profiles of dominant FFN sub-updates when a prediction has already saturated.

  9. Knowl 9 — Early Exit Efficiency and Accuracy Benchmark on WikiLM

    data/table

    Evaluation of early exit methods on WikiLM over WikiText-103 validation data compares the cluster-overlap rule based strictly on dominant sub-updates against layer-wise binary logistic regression classifiers trained directly on hidden state representations xℓx^\ell, FFN output vectors oℓo^\ell, and residual representations x~ℓ=xℓ+oℓ\tilde{x}^\ell = x^\ell + o^\ell.

    Method Prediction Accuracy (%) Saved Layers (%)
    Binary classifiers using xℓx^\ell 94.4±6.494.4 \pm 6.4 18.8%±0.418.8\% \pm 0.4 (3.0±0.43.0 \pm 0.4 layers)
    Binary classifiers using oℓo^\ell 92.9±5.492.9 \pm 5.4 19.4%±0.319.4\% \pm 0.3 (3.1±0.33.1 \pm 0.3 layers)
    Binary classifiers using x~ℓ\tilde{x}^\ell 94.4±6.494.4 \pm 6.4 18.8%±0.418.8\% \pm 0.4 (3.0±0.43.0 \pm 0.4 layers)
    Sub-updates cluster rule 94.1±1.494.1 \pm 1.4 20.0%±0.220.0\% \pm 0.2 (3.2±0.33.2 \pm 0.3 layers)

    Prediction accuracy measures the fraction of early-exited examples that produce identical final predictions to the full 16-layer model. The training-free sub-updates rule achieves 94.1% accuracy while bypassing 20.0% of total network layers (saving an average of 3.2 layers out of 16), performing on par with supervised logistic regression classifiers trained on dense continuous representations.

  10. Knowl 10 — Preservation of Concept Projections Under Layer Normalization

    empirical result

    In transformer architectures incorporating post-FFN layer normalization (such as GPT-2), projecting unnormalized value vectors EviℓE v_i^\ell preserves the underlying concept rankings compared to projecting normalized vectors E⋅LayerNorm(viℓ)E \cdot \text{LayerNorm}(v_i^\ell).

    Measuring the similarity between the top-30 scoring vocabulary tokens of unnormalized projections EviℓE v_i^\ell and normalized projections E⋅LayerNorm(viℓ)E \cdot \text{LayerNorm}(v_i^\ell) using Intersection over Union (IoU) across all 24 layers of GPT-2 yields an average overlap of 64.5%. In comparison, evaluating IoU between unnormalized value vector projections and random vectors sampled from a matching Gaussian distribution yields an average overlap of approximately 0%. This demonstrates that direct linear projection EviℓE v_i^\ell faithfully extracts the semantic and syntactic concepts promoted by FFN value vectors without requiring explicit layer normalization modeling.

Coverage note — None omitted; all primary conceptual formulations, empirical analyses (concept interpretability, promotion mechanism, dominant sub-update contributions, functional vector clustering, layer norm effects), and applications (toxic language suppression and early exit prediction) are fully captured.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  2. 2.Alexei Baevski and Michael Auli. 2019. Adaptive input representations for neural language modeling. In International Conference on Learning Representations (ICLR).
  3. 3.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT).
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  5. 5.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy. Association for Computational Linguistics.
  6. 6.Jeff Da, Ronan Le Bras, Ximing Lu, Yejin Choi, and Antoine Bosselut. 2021. Analyzing commonsense emergence in few-shot knowledge models. In 3rd Conference on Automated Knowledge Base Construction.
  7. 7.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland. Association for Computational Linguistics.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Association for Computational Linguistics (NAACL), pages 4171–4186, Minneapolis, Minnesota.
  9. 9.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2020. Depth-adaptive transformer. In International Conference on Learning Representations (ICLR).
  10. 10.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread. Https://transformer-circuits.pub/2021/framework/index.html.
  11. 11.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  12. 12.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the conference on computer vision and pattern recognition (CVPR).
  14. 14.Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415.
  15. 15.Arthur E Hoerl and Robert W Kennard. 1970. Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67.
  16. 16.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020. Dynabert: Dynamic bert with adaptive width and depth. Advances in Neural Information Processing Systems (NeurIPS).
  17. 17.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  18. 18.Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, and Yejin Choi. 2022. Prompt waywardness: The curious case of discretized interpretation of continuous prompts. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3631–3643, Seattle, United States. Association for Computational Linguistics.
  19. 19.Lei Li, Yankai Lin, Deli Chen, Shuhuai Ren, Peng Li, Jie Zhou, and Xu Sun. 2021. CascadeBERT: Accelerating inference of pre-trained language models via calibrated complete models cascade. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 475–486, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  20. 20.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  21. 21.Kris McGuffie and Alex Newhouse. 2020. The radicalization risks of gpt-3 and advanced neural language models. arXiv preprint arXiv:2009.06807.
  22. 22.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual knowledge in gpt. arXiv preprint arXiv:2202.05262.
  23. 23.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. International Conference on Learning Representations (ICLR).
  24. 24.Daniel Müllner. 2011. Modern hierarchical, agglomerative clustering algorithms. arXiv preprint arXiv:1109.2378.
  25. 25.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  26. 26.Naomi Saphra and Adam Lopez. 2019. Understanding learning dynamics of language models with SVCCA. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3257–3267, Minneapolis, Minnesota. Association for Computational Linguistics.
  27. 27.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  28. 28.Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay. 2021. Consistent accelerated inference via confident adaptive transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4962–4979, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  29. 29.Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020a. Green AI. Communications of the ACM, 63(12):54–63.
  30. 30.Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. 2020b. The right tool for the job: Matching model and instance complexities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6640–6651, Online. Association for Computational Linguistics.
  31. 31.Noam M. Shazeer. 2020. Glu variants improve transformer. ArXiv, abs/2002.05202.
  32. 32.S. Sukhbaatar, J. Weston, and R. Fergus. 2015. End-to-end memory networks. In Advances in Neural Information Processing Systems (NIPS).
  33. 33.Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019. Augmenting self-attention with persistent memory. arXiv preprint arXiv:1907.01470.
  34. 34.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy. Association for Computational Linguistics.
  35. 35.Robert Tibshirani. 1996. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58(1):267–288.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NIPS), pages 5998–6008.
  37. 37.Elena Voita, Rico Sennrich, and Ivan Titov. 2019. The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4396–4406, Hong Kong, China. Association for Computational Linguistics.
  38. 38.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2153–2162, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2020. BERTnesia: Investigating the capture and forgetting of knowledge in BERT. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 174–183, Online. Association for Computational Linguistics.
  40. 40.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022. What language model architecture and pretraining objective works best for zero-shot generalization? In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 22964–22984. PMLR.
  41. 41.Ji Xin, Rodrigo Nogueira, Yaoliang Yu, and Jimmy Lin. 2020. Early exiting BERT for efficient document ranking. In Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, pages 83–88, Online. Association for Computational Linguistics.
  42. 42.Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. BERxiT: Early exiting for BERT with better fine-tuning and extension to regression. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 91–104, Online. Association for Computational Linguistics.
  43. 43.Jingjing Xu, Wangchunshu Zhou, Zhiyi Fu, Hao Zhou, and Lei Li. 2021. A survey on green deep learning. arXiv preprint arXiv:2111.05193.
  44. 44.Yilin Yang, Longyue Wang, Shuming Shi, Prasad Tadepalli, Stefan Lee, and Zhaopeng Tu. 2020. On the sub-layer functionalities of transformer decoder. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4799–4811, Online. Association for Computational Linguistics.
  45. 45.Yunzhi Yao, Shaohan Huang, Ningyu Zhang, Li Dong, Furu Wei, and Huajun Chen. 2022. Kformer: Knowledge injection in transformer feed-forward layers. arXiv preprint arXiv:2201.05742.

Citation

MLA
Geva, M., et al. “Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 30–45, https://doi.org/10.18653/v1/2022.emnlp-main.3.
APA
Geva, M., Caciularu, A., Wang, K., & Goldberg, Y. (2022). Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 30–45. https://doi.org/10.18653/v1/2022.emnlp-main.3
Chicago
Geva, M., A. Caciularu, K. Wang, and Y. Goldberg. 2022. “Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 30–45. https://doi.org/10.18653/v1/2022.emnlp-main.3.
Harvard
Geva, M. et al. (2022) “Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 30–45. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.3.
Vancouver
1. Geva M, Caciularu A, Wang K, Goldberg Y (2022) Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 30–45

BibTeX

@inproceedings{geva-etal-2022-transformer,
    title = "Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space",
    author = "Geva, Mor  and
      Caciularu, Avi  and
      Wang, Kevin  and
      Goldberg, Yoav",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.3/",
    doi = "10.18653/v1/2022.emnlp-main.3",
    pages = "30--45"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/