GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers

Ali ModarressiMohsen FayyazYadollah YaghoobzadehMohammad Taher Pilehvar

article2022NAACL63 citations

Proposes GlobEnc, a layer-aggregation framework that measures global token attribution by accounting for full Transformer encoder components, including layer normalizations and residual connections, to achieve higher correlation with gradient-based saliency scores than attention-only baselines.

Listen

Modern natural language processing relies heavily on Transformer models, yet interpreting how these complex architectures arrive at specific decisions remains an active challenge. Standard interpretability techniques often rely on raw attention weights or focus exclusively on the self-attention sub-layer, which produces misleading explanations. While gradient-based methods provide more reliable importance scores, they are computationally intensive and slow. The article introduces and evaluates GlobEnc, a novel method designed to quantify how much each input word contributes to a model's final output by incorporating almost all components within the entire encoder block across all layers.

The authors evaluated GlobEnc using fine-tuned BERT-base, BERT-large, and ELECTRA-base models across three benchmark text classification tasks: sentiment analysis (SST2), natural language inference (MNLI), and hate speech detection (HateXplain). The approach accounts for vector magnitudes (norms), residual connections, and both layer normalizations within the encoder layer, then aggregates these layer-level contributions across the entire network using an attention rollout technique. The resulting global attribution scores were benchmarked against gradient-based saliency scores to measure explanation faithfulness.

The analysis revealed that GlobEnc consistently outperforms existing attention-based attribution methods, achieving high Spearman rank correlations with gradient saliency across all datasets (reaching 0.72 to 0.78 on BERT-base and up to 0.83 on BERT-large), while raw attention weights showed negligible or negative correlations (ranging from -0.11 to 0.12 on BERT-base). The authors found that accounting for vector norms and exact residual connections is essential, as models naturally preserve individual word representations far more than mixing them. Crucially, the study discovered that including only the first layer normalization severely degrades attribution accuracy due to outlier weights, but incorporating both layer normalizations restores high fidelity because their outlier effects counteract one another. Furthermore, multi-layer aggregation proved vital, showing that models identify key decision-driving words within the first few layers before stabilizing.

These findings demonstrate that organizations can achieve highly faithful model explanations without the prohibitive computational costs of gradient-based techniques, as computing the proposed encoder maps takes seconds compared to hours for gradient alternatives. This provides a practical, scalable mechanism for auditing model behavior, ensuring regulatory compliance, and managing performance risks in high-stakes deployments.

Teams seeking to interpret Transformer architectures should transition away from raw attention weights or isolated sub-layer analyses in favor of full encoder layer aggregation. Future efforts should focus on validating GlobEnc across generative language models and broader operational datasets before fully standardizing its use in production monitoring.

The main limitation of this approach is its necessary omission of the non-linear feed-forward sub-layer, which cannot be linearly decomposed into token contributions. Nevertheless, confidence in the reported results is high, as the strong correlation gains remain robust across diverse model architectures, sizes, and evaluation tasks.

arXiv: 2205.03286mohsenfayyaz/GlobEnc
  • Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). This paper establishes the foundational post hoc aggregation methods (attention rollout and attention flow) across Transformer layers, highlighting the failure of raw attention weights and motivating GlobEnc's whole-encoder attribution approach.
  • Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This work demonstrates that raw attention weights do not provide faithful explanations, laying the conceptual groundwork for developing attribution methods that incorporate other layer components.
  • Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This paper reveals the functional importance of feed-forward networks in Transformers, underpinning GlobEnc's premise that non-attention components in the encoder block must be integrated into attribution analysis.
  • Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). This foundational study analyzes individual attention head behaviors in BERT, providing key empirical context on how Transformers route information across layers.
  • Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). This work introduces Integrated Gradients and axiomatic feature attribution, which serve as standard gradient-based saliency baselines used to evaluate token attribution faithfuless in GlobEnc.
  • Paper: Methods for interpreting and understanding deep neural networks, Grégoire Montavon et al. (2018). This comprehensive overview covers decomposition-based explanation techniques like Layer-wise Relevance Propagation that inspire layer-by-layer attribution aggregation in neural networks.
Cover for GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers

Abstract

There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies have shown that integrating other components can yield more accurate explanations. This paper introduces a novel token attribution analysis method that incorporates all the components in the encoder block and aggregates this throughout layers. Through extensive quantitative and qualitative experiments, we demonstrate that our method can produce faithful and meaningful global token attributions. Our experiments reveal that incorporating almost every encoder component results in increasingly more accurate analysis in both local (single layer) and global (the whole model) settings. Our global attribution analysis significantly outperforms previous methods on various tasks regarding correlation with gradient-based saliency scores. Our code is freely available at https://github.com/mohsenfayyaz/GlobEnc.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background
  • 3 Methodology
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Analysis Methods
  • 4.3 Gradient-based Methods for Faithfulness Analysis
  • 4.3.1 Saliency
  • 4.3.2 HTA x Inputs
  • 4.4 Setup
  • 4.5 Results
  • 4.5.1 On the role of vector norms
  • 4.5.2 On the role of residual connections
  • 4.5.3 On the role of layer normalization
  • 4.5.4 On the role of aggregation
  • 4.5.5 Qualitative analysis
  • 5 Related Work
  • 6 Conclusions
  • References
  • A Appendix
  • A.1 LN Formulation
  • A.2 More Models
  • A.3 More Examples

Knowls

  1. Knowl 1 — GlobEnc Encoder-Layer Token Attribution Decomposition

    model/method

    In a Transformer encoder layer, let x1,…,xn∈Rdx_1, \dots, x_n \in \mathbb{R}^d be the input token representations. The self-attention block computes transformed vectors ziz_i, which are added to the first residual connection (RES#1) to give zi+=zi+xiz_i^+ = z_i + x_i, and normalized by the first layer normalization (LN#1) to produce z~i=LN(zi+)\tilde{z}_i = \text{LN}(z_i^+). This representation is passed through a feed-forward network (FFN), added to a second residual connection (RES#2) yielding z~i+=FFN(z~i)+z~i\tilde{z}_i^+ = \text{FFN}(\tilde{z}_i) + \tilde{z}_i, and normalized by the second layer normalization (LN#2) to obtain the encoder layer output representation x~i=LN(z~i+)\tilde{x}_i = \text{LN}(\tilde{z}_i^+).

    The decomposition of z~i\tilde{z}_i into contributions from input token jj is given by: z~i←j=gzi+(∑h=1Hαi,jhfh(xj)+1[i=j]xi)\tilde{z}_{i \leftarrow j} = g_{z_i^+}\left(\sum_{h=1}^H \alpha_{i,j}^h f^h(x_j) + \mathbf{1}[i = j] x_i\right) where HH is the number of attention heads, αi,jh\alpha_{i,j}^h is the raw attention weight from token ii to token jj in head hh, fh(xj)=vh(xj)WOhf^h(x_j) = v^h(x_j) W_O^h represents the projected value vector with head-specific projection slice WOhW_O^h, and gz(u):=u−m(u)s(z)⊙γg_{z}(u) := \frac{u - m(u)}{s(z)} \odot \gamma decomposes layer normalization with element-wise mean m(u)=1d∑ku(k)m(u) = \frac{1}{d} \sum_k u^{(k)}, standard deviation s(z)=1d∑k(m(z)−z(k)+ϵ)2s(z) = \sqrt{\frac{1}{d} \sum_k (m(z) - z^{(k)} + \epsilon)^2}, and learnable scale parameter γ∈Rd\gamma \in \mathbb{R}^d (omitting the bias β\beta).

    Because the non-linear activation function in the FFN prevents an exact linear decomposition of FFN outputs across tokens, the direct contribution of token jj through the FFN path is omitted while retaining RES#2 and preserving the FFN's influence through the standard deviation s(z~i+)s(\tilde{z}_i^+) of LN#2. The resulting token attribution vector x~i←j\tilde{x}_{i \leftarrow j} from input token jj to output token ii of the encoder layer is approximated as: x~i←j≈gz~i+(z~i←j)=z~i←j−m(z~i←j)s(z~i+)⊙γ\tilde{x}_{i \leftarrow j} \approx g_{\tilde{z}_i^+}(\tilde{z}_{i \leftarrow j}) = \frac{\tilde{z}_{i \leftarrow j} - m(\tilde{z}_{i \leftarrow j})}{s(\tilde{z}_i^+)} \odot \gamma The local encoder attribution matrix NENC∈Rn×n\mathcal{N}_{\text{ENC}} \in \mathbb{R}^{n \times n} is defined by (NENC)i,j:=∥x~i←j∥(\mathcal{N}_{\text{ENC}})_{i,j} := \|\tilde{x}_{i \leftarrow j}\|, measuring the magnitude of token jj's contribution to token ii across the encoder layer.

  2. Knowl 2 — GlobEnc Global Token Attribution Aggregation via Modified Attention Rollout

    algorithm

    GlobEnc aggregates layerwise attribution matrices NENC(ℓ)\mathcal{N}_{\text{ENC}}^{(\ell)} across all LL layers of a Transformer encoder to produce a global token attribution vector for the sequence classification token [CLS][\text{CLS}]. Unlike standard attention rollout—which adds an artificial fixed identity matrix 0.5I0.5 I to account for residual connections—GlobEnc operates directly on normalized attribution matrices because NENC\mathcal{N}_{\text{ENC}} already accounts for residual connections and layer normalizations.

    Input: Sequence of layerwise attribution matrices NENC(ℓ)∈Rn×n\mathcal{N}_{\text{ENC}}^{(\ell)} \in \mathbb{R}^{n \times n} for layers ℓ=1,…,L\ell = 1, \dots, L.
    Output: Global input token attribution vector a∈Rna \in \mathbb{R}^n for the [CLS][\text{CLS}] token.
    for ℓ=1\ell = 1 to LL:
        for i=1i = 1 to nn:
            Si←∑k=1n(NENC(ℓ))i,kS_i \leftarrow \sum_{k=1}^n (\mathcal{N}_{\text{ENC}}^{(\ell)})_{i,k}
            for j=1j = 1 to nn:
                (M(ℓ))i,j←(NENC(ℓ))i,j/Si(M^{(\ell)})_{i,j} \leftarrow (\mathcal{N}_{\text{ENC}}^{(\ell)})_{i,j} / S_i
        if ℓ==1\ell == 1:
            A~1←M(1)\tilde{\mathcal{A}}_1 \leftarrow M^{(1)}
        else:
            A~ℓ←M(ℓ)A~ℓ−1\tilde{\mathcal{A}}_\ell \leftarrow M^{(\ell)} \tilde{\mathcal{A}}_{\ell-1}
    a←row1(A~L)a \leftarrow \text{row}_1(\tilde{\mathcal{A}}_L)
    return aa

    The resulting accumulated matrix A~L∈Rn×n\tilde{\mathcal{A}}_L \in \mathbb{R}^{n \times n} represents the aggregated information flow from input tokens to final layer representations. The first row (i=1i=1, corresponding to [CLS][\text{CLS}]) defines the importance of each input token to the classification decision.

  3. Knowl 3 — Counteracting Outlier Weights in Transformer Layer Normalizations

    empirical result

    Incorporating only the attention block's first layer normalization (LN#1) into norm-based attribution (denoted NRESLN\mathcal{N}_{\text{RESLN}}) causes a sharp drop in correlation with gradient-based saliency compared to analyzing attention and residual connections alone (NRES\mathcal{N}_{\text{RES}}). On BERT-base fine-tuned on SST-2, the Spearman rank correlation drops from 0.730.73 for NRES\mathcal{N}_{\text{RES}} to −0.21-0.21 for NRESLN\mathcal{N}_{\text{RESLN}}.

    This degradation is caused by the learned outlier weights in the layer normalization parameters γ\gamma (dimensions where weights deviate by at least 3σ3\sigma from the layer mean). The outlier weights of LN#1 and LN#2 are strongly negatively correlated with each other across dimensions, exhibiting negative Pearson correlation coefficients across layers (reaching approximately −0.90-0.90 in layer 11). As a result, the distortion introduced by outlier dimensions in LN#1 is cancelled out when LN#2 is also included in the analysis (NENC\mathcal{N}_{\text{ENC}}), which restores the correlation to 0.770.77. Consequently, the two layer normalizations in a Transformer encoder must be modeled together rather than in isolation.

  4. Knowl 4 — Global Attribution Faithfulness on Fine-Tuned BERT-base

    data/table

    The faithfulness of token attribution methods is evaluated by measuring the Spearman rank correlation between the [CLS][\text{CLS}] token's aggregated attribution scores at the final layer and gradient ×\times input saliency scores, defined as Saliencyi=∥∂yc∂ei0⊙ei0∥2\text{Saliency}_i = \|\frac{\partial y_c}{\partial e_i^0} \odot e_i^0\|_2 for true class score ycy_c and input embedding ei0e_i^0. Evaluations are performed on BERT-base fine-tuned on SST-2, MNLI, and HateXplain datasets using rollout aggregation across layers.

    Method SST2 MNLI HateXplain
    Weight-based (WW) −0.11±0.26-0.11 \pm 0.26 −0.06±0.22-0.06 \pm 0.22 0.12±0.260.12 \pm 0.26
    w/ Fixed Residual (WFIXEDRESW_{\text{FIXEDRES}}) −0.24±0.26-0.24 \pm 0.26 −0.05±0.26-0.05 \pm 0.26 0.13±0.280.13 \pm 0.28
    w/ Residual (WRESW_{\text{RES}}) 0.19±0.260.19 \pm 0.26 0.27±0.250.27 \pm 0.25 0.53±0.240.53 \pm 0.24
    Norm-based (N\mathcal{N}) 0.44±0.200.44 \pm 0.20 0.47±0.160.47 \pm 0.16 0.43±0.220.43 \pm 0.22
    w/ Fixed Residual (NFIXEDRES\mathcal{N}_{\text{FIXEDRES}}) 0.48±0.200.48 \pm 0.20 0.55±0.160.55 \pm 0.16 0.48±0.220.48 \pm 0.22
    w/ Residual (NRES\mathcal{N}_{\text{RES}}) 0.73±0.130.73 \pm 0.13 0.75±0.100.75 \pm 0.10 0.66±0.170.66 \pm 0.17
    w/ Residual + Layer Norm 1 (NRESLN\mathcal{N}_{\text{RESLN}}) −0.21±0.26-0.21 \pm 0.26 −0.06±0.26-0.06 \pm 0.26 0.08±0.280.08 \pm 0.28
    w/ GlobEnc: [Residual + Layer Norm 1, 2] (NENC\mathcal{N}_{\text{ENC}}) 0.77±0.12\mathbf{0.77 \pm 0.12} 0.78±0.09\mathbf{0.78 \pm 0.09} 0.72±0.17\mathbf{0.72 \pm 0.17}

    The results show that:

    1. Norm-based formulations outperform weight-based equivalents across all benchmarks.
    2. Dynamic residual mixing derived from vector norms (WRESW_{\text{RES}}, NRES\mathcal{N}_{\text{RES}}) outperforms uniform fixed residual assumptions (WFIXEDRESW_{\text{FIXEDRES}}, NFIXEDRES\mathcal{N}_{\text{FIXEDRES}}).
    3. GlobEnc (NENC\mathcal{N}_{\text{ENC}}) achieves the highest correlation on all three classification tasks.
  5. Knowl 5 — Input-Scaled Hidden Token Attribution (HTA x Inputs)

    equation

    To quantify the sensitivity and information mixing between consecutive layers in a Transformer network, Hidden Token Attribution (HTA) is scaled by the hidden embedding vector values. For hidden embedding ejℓ−1∈Rde_j^{\ell-1} \in \mathbb{R}^d of token jj at layer ℓ−1\ell-1 and hidden embedding eiℓ∈Rde_i^\ell \in \mathbb{R}^d of token ii at layer ℓ\ell, the input-scaled HTA attribution score ci←jℓc_{i \leftarrow j}^\ell is defined as:

    ci←jℓ=∥∂eiℓ∂ejℓ−1⊙ejℓ−1∥Fc_{i \leftarrow j}^\ell = \left\| \frac{\partial e_i^\ell}{\partial e_j^{\ell-1}} \odot e_j^{\ell-1} \right\|_F

    where ∂eiℓ∂ejℓ−1∈Rd×d\frac{\partial e_i^\ell}{\partial e_j^{\ell-1}} \in \mathbb{R}^{d \times d} is the Jacobian matrix, ⊙\odot represents column-wise element-wise multiplication broadcasting ejℓ−1e_j^{\ell-1}, and ∥⋅∥F\|\cdot\|_F denotes the Frobenius norm.

  6. Knowl 6 — Context-Mixing Ratio Correction for Weight-Based Attention Rollout

    model/method

    Standard attention rollout sets a fixed residual mixing ratio ri≈0.5r_i \approx 0.5, which incorrectly assumes that Transformer layers mix contextual information and preserve self-representations in equal proportions. To separate the contribution of accurate residual proportioning from vector norm calculations, a dynamic context-mixing ratio r^i\hat{r}_i is derived from the norm-based attribution vectors x~i←j\tilde{x}_{i \leftarrow j}:

    r^i=∥∑j=1,j≠inx~i←j∥∥∑j=1,j≠inx~i←j∥+∥x~i←i∥\hat{r}_i = \frac{\left\| \sum_{j=1, j \neq i}^n \tilde{x}_{i \leftarrow j} \right\|}{\left\| \sum_{j=1, j \neq i}^n \tilde{x}_{i \leftarrow j} \right\| + \|\tilde{x}_{i \leftarrow i}\|}

    To construct the corrected weight-based attribution matrix WRESW_{\text{RES}}, the head-averaged attention matrix Aˉℓ\bar{A}_\ell is modified by zeroing its diagonal elements, renormalizing each row, and interpolating with the identity matrix II using r^i\hat{r}_i:

    Aℓ′=(I−diag(Aˉℓ))−1(Aˉℓ−diag(Aˉℓ))A'_\ell = (I - \text{diag}(\bar{A}_\ell))^{-1} (\bar{A}_\ell - \text{diag}(\bar{A}_\ell)) WRES:=diag(r^1,…,r^n)Aℓ′+diag(1−r^1,…,1−r^n)IW_{\text{RES}} := \text{diag}(\hat{r}_1, \dots, \hat{r}_n) A'_\ell + \text{diag}(1 - \hat{r}_1, \dots, 1 - \hat{r}_n) I

    where diag(v)\text{diag}(v) creates a diagonal matrix from vector vv. Incorporating this dynamic ratio raises the correlation of weight-based rollout from −0.24-0.24 (WFIXEDRESW_{\text{FIXEDRES}}) to 0.190.19 (WRESW_{\text{RES}}) on SST-2.

  7. Knowl 7 — Layer-wise vs. Aggregated Attribution Performance

    data/table

    Spearman's rank correlation between saliency scores and attribution scores is compared for individual single layers versus multi-layer rollout aggregation on BERT-base fine-tuned on SST-2.

    Method Layer 1 Layer 6 Layer 12 Max Layer
    Individual Single Layer
    N\mathcal{N} −0.50±0.18-0.50 \pm 0.18 +0.28±0.23+0.28 \pm 0.23 +0.40±0.21+0.40 \pm 0.21 +0.41±0.21+0.41 \pm 0.21
    NRES\mathcal{N}_{\text{RES}} −0.48±0.18-0.48 \pm 0.18 +0.29±0.24+0.29 \pm 0.24 +0.41±0.19+0.41 \pm 0.19 +0.41±0.19+0.41 \pm 0.19
    NENC\mathcal{N}_{\text{ENC}} −0.47±0.18-0.47 \pm 0.18 +0.29±0.24+0.29 \pm 0.24 +0.41±0.19+0.41 \pm 0.19 +0.41±0.19+0.41 \pm 0.19
    Aggregated via Rollout
    N\mathcal{N} −0.50±0.18-0.50 \pm 0.18 +0.44±0.20+0.44 \pm 0.20 +0.44±0.20+0.44 \pm 0.20 +0.44±0.20+0.44 \pm 0.20
    NRES\mathcal{N}_{\text{RES}} −0.48±0.18-0.48 \pm 0.18 +0.70±0.14+0.70 \pm 0.14 +0.73±0.13+0.73 \pm 0.13 +0.73±0.13+0.73 \pm 0.13
    NENC\mathcal{N}_{\text{ENC}} (GlobEnc) −0.47±0.18-0.47 \pm 0.18 +0.74±0.14\mathbf{+0.74 \pm 0.14} +0.77±0.12\mathbf{+0.77 \pm 0.12} +0.78±0.12\mathbf{+0.78 \pm 0.12}

    Attributions extracted from individual single layers peak at a correlation of +0.41+0.41, whereas aggregating layerwise attributions across layers via rollout increases the correlation to +0.78+0.78 for NENC\mathcal{N}_{\text{ENC}}, demonstrating that cross-layer aggregation is necessary for accurate global token attribution.

  8. Knowl 8 — GlobEnc Attribution Performance on BERT-large and ELECTRA-base

    data/table

    The performance of rollout-aggregated attribution methods was evaluated across different model scales and pre-training objectives using 24-layer BERT-large and 12-layer ELECTRA-base (trained with replaced token detection). Faithfulness is reported as Spearman's rank correlation with gradient ×\times input saliency scores across SST-2, MNLI, and HateXplain validation sets.

    Method SST2 MNLI HateXplain
    BERT-large
    Weight-based (WW) −0.38±0.16-0.38 \pm 0.16 −0.61±0.14-0.61 \pm 0.14 −0.41±0.25-0.41 \pm 0.25
    w/ Fixed Residual (WFIXEDRESW_{\text{FIXEDRES}}) −0.25±0.19-0.25 \pm 0.19 −0.48±0.19-0.48 \pm 0.19 −0.21±0.30-0.21 \pm 0.30
    w/ Residual (WRESW_{\text{RES}}) −0.10±0.21-0.10 \pm 0.21 0.33±0.230.33 \pm 0.23 0.09±0.300.09 \pm 0.30
    Norm-based (N\mathcal{N}) 0.44±0.240.44 \pm 0.24 0.13±0.270.13 \pm 0.27 0.48±0.250.48 \pm 0.25
    w/ Fixed Residual (NFIXEDRES\mathcal{N}_{\text{FIXEDRES}}) 0.49±0.240.49 \pm 0.24 0.26±0.250.26 \pm 0.25 0.49±0.300.49 \pm 0.30
    w/ Residual (NRES\mathcal{N}_{\text{RES}}) 0.77±0.110.77 \pm 0.11 0.66±0.120.66 \pm 0.12 0.73±0.160.73 \pm 0.16
    w/ Residual + Layer Norm 1 (NRESLN\mathcal{N}_{\text{RESLN}}) −0.07±0.23-0.07 \pm 0.23 −0.35±0.24-0.35 \pm 0.24 0.06±0.320.06 \pm 0.32
    w/ GlobEnc (NENC\mathcal{N}_{\text{ENC}}) 0.83±0.08\mathbf{0.83 \pm 0.08} 0.77±0.09\mathbf{0.77 \pm 0.09} 0.76±0.17\mathbf{0.76 \pm 0.17}
    ELECTRA-base
    Weight-based (WW) −0.37±0.19-0.37 \pm 0.19 −0.31±0.22-0.31 \pm 0.22 0.02±0.290.02 \pm 0.29
    w/ Fixed Residual (WFIXEDRESW_{\text{FIXEDRES}}) −0.37±0.19-0.37 \pm 0.19 −0.24±0.23-0.24 \pm 0.23 0.01±0.290.01 \pm 0.29
    w/ Residual (WRESW_{\text{RES}}) −0.10±0.22-0.10 \pm 0.22 0.08±0.250.08 \pm 0.25 0.20±0.270.20 \pm 0.27
    Norm-based (N\mathcal{N}) 0.18±0.210.18 \pm 0.21 0.12±0.210.12 \pm 0.21 0.21±0.260.21 \pm 0.26
    w/ Fixed Residual (NFIXEDRES\mathcal{N}_{\text{FIXEDRES}}) 0.23±0.220.23 \pm 0.22 0.32±0.230.32 \pm 0.23 0.28±0.260.28 \pm 0.26
    w/ Residual (NRES\mathcal{N}_{\text{RES}}) 0.54±0.170.54 \pm 0.17 0.54±0.140.54 \pm 0.14 0.44±0.210.44 \pm 0.21
    w/ Residual + Layer Norm 1 (NRESLN\mathcal{N}_{\text{RESLN}}) −0.24±0.23-0.24 \pm 0.23 −0.16±0.24-0.16 \pm 0.24 −0.07±0.28-0.07 \pm 0.28
    w/ GlobEnc (NENC\mathcal{N}_{\text{ENC}}) 0.64±0.15\mathbf{0.64 \pm 0.15} 0.68±0.12\mathbf{0.68 \pm 0.12} 0.47±0.22\mathbf{0.47 \pm 0.22}

    GlobEnc achieves superior rank correlations across both deep 24-layer models (up to 0.830.83 on SST-2) and discriminator-pretrained architectures, confirming the general applicability of the full-encoder norm aggregation approach.

  9. Knowl 9 — Omission of Non-Linear Feed-Forward Network in GlobEnc Token Decomposition

    limitation

    In a Transformer encoder, the intermediate representation z~i\tilde{z}_i after the attention block is processed by a two-layer feed-forward network with a non-linear activation function: FFN(z~i)=max⁡(0,z~iW1+b1)W2+b2\text{FFN}(\tilde{z}_i) = \max(0, \tilde{z}_i W_1 + b_1) W_2 + b_2 Because the activation function is non-linear, a linear decomposition of the FFN output into token-specific contribution terms ∑j=1nFFN(z~i←j)\sum_{j=1}^n \text{FFN}(\tilde{z}_{i \leftarrow j}) is not mathematically possible.

    To preserve an exact closed-form attribution calculation, GlobEnc omits the token-level transformation through the FFN, approximating the input to the second layer normalization using only the residual representation z~i←j\tilde{z}_{i \leftarrow j}. The FFN still exerts a partial influence on the resulting score x~i←j=z~i←j−m(z~i←j)s(z~i+)⊙γ\tilde{x}_{i \leftarrow j} = \frac{\tilde{z}_{i \leftarrow j} - m(\tilde{z}_{i \leftarrow j})}{s(\tilde{z}_i^+)} \odot \gamma because the vector z~i+=FFN(z~i)+z~i\tilde{z}_i^+ = \text{FFN}(\tilde{z}_i) + \tilde{z}_i determines the normalization standard deviation s(z~i+)s(\tilde{z}_i^+).

Coverage note — None was omitted; all contributed methodology, formulas, empirical tables, and analytical findings are fully captured.

References

  1. 1.Samira Abnar and Willem Zuidema. 2020. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, Online. Association for Computational Linguistics.
  2. 2.Jay Alammar. 2018. The illustrated transformer [blog post].
  3. 3.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. A diagnostic study of explainability techniques for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3256–3274, Online. Association for Computational Linguistics.
  4. 4.Jasmijn Bastings and Katja Filippova. 2020. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 149–155, Online. Association for Computational Linguistics.
  5. 5.Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020. On identifiability in transformers. In International Conference on Learning Representations.
  6. 6.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy. Association for Computational Linguistics.
  7. 7.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pretraining text encoders as discriminators rather than generators. In ICLR.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2021. Not all models localize linguistic knowledge in the same place: A layer-wise probing on BERToids’ representations. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 375–388, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman. 2019. Do attention heads in BERT track syntactic dependencies? CoRR, abs/1911.12246.
  11. 11.Sarthak Jain and Byron C. Wallace. 2019. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3543–3556, Minneapolis, Minnesota.
  12. 12.Pieter-Jan Kindermans, Kristof Schütt, Klaus-Robert Müller, and Sven Dähne. 2016. Investigating the influence of noise and distractors on the interpretation of neural networks.
  13. 13.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7057–7075, Online. Association for Computational Linguistics.
  14. 14.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2021. Incorporating Residual and Normalization Layers into Analysis of Masked Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4547–4568, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky. 2021. BERT busters: Outlier dimensions that disrupt transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3392–3405, Online. Association for Computational Linguistics.
  16. 16.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4365–4374, Hong Kong, China.
  17. 17.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016. Visualizing and understanding neural models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681–691, San Diego, California. Association for Computational Linguistics.
  18. 18.Ziyang Luo, Artur Kulmizev, and Xiaoxi Mao. 2021. Positional artefacts propagate through masked language model embeddings. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5312–5327, Online. Association for Computational Linguistics.
  19. 19.Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. Hatexplain: A benchmark dataset for explainable hate speech detection. Proceedings of the AAAI Conference on Artificial Intelligence, 35(17):14867–14875.
  20. 20.Hosein Mohebbi, Ali Modarressi, and Mohammad Taher Pilehvar. 2021. Exploring the role of BERT token representations to explain sentence probing results. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 792–806, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  21. 21.Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2021. Telling BERT’s full story: from local attention to global aggregation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 105–124, Online. Association for Computational Linguistics.
  22. 22.Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019. Visualizing and measuring the geometry of bert. In Advances in Neural Information Processing Systems, pages 8594–8603.
  23. 23.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. In NeurIPS EMC2 Workshop.
  24. 24.Sofia Serrano and Noah A. Smith. 2019. Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2931–2951, Florence, Italy.
  25. 25.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Deep inside convolutional networks: Visualising image classification models and saliency maps. CoRR, abs/1312.6034.
  26. 26.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  27. 27.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  28. 28.Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20, Hong Kong, China. Association for Computational Linguistics.
  29. 29.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  30. 30.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  31. 31.Zhengxuan Wu and Desmond C. Ong. 2021. On explaining your explanations of BERT: an empirical study with sequence classification. CoRR, abs/2101.00196.
  32. 32.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020. Bert loses patience: Fast and robust inference with early exit. In Advances in Neural Information Processing Systems, volume 33, pages 18330–18341. Curran Associates, Inc.

Citation

MLA
Modarressi, A., et al. “GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 258–71, https://doi.org/10.18653/v1/2022.naacl-main.19.
APA
Modarressi, A., Fayyaz, M., Yaghoobzadeh, Y., & Pilehvar, M. T. (2022). GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 258–271. https://doi.org/10.18653/v1/2022.naacl-main.19
Chicago
Modarressi, A., M. Fayyaz, Y. Yaghoobzadeh, and M. T. Pilehvar. 2022. “GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 258–71. https://doi.org/10.18653/v1/2022.naacl-main.19.
Harvard
Modarressi, A. et al. (2022) “GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 258–271. Available at: https://doi.org/10.18653/v1/2022.naacl-main.19.
Vancouver
1. Modarressi A, Fayyaz M, Yaghoobzadeh Y, Pilehvar MT (2022) GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 258–271

BibTeX

@inproceedings{modarressi-etal-2022-globenc,
    title = "{G}lob{E}nc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers",
    author = "Modarressi, Ali  and
      Fayyaz, Mohsen  and
      Yaghoobzadeh, Yadollah  and
      Pilehvar, Mohammad Taher",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.19/",
    doi = "10.18653/v1/2022.naacl-main.19",
    pages = "258--271"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/